Multiple voice services

Networked microphone devices enhance media playback systems by integrating voice control capabilities, enabling synchronized and dynamic control of multiple playback devices and services across zones, addressing limitations in existing systems.

JP7698681B2Active Publication Date: 2025-06-25SONOS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023144387
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-08-05
Filing Date
2023-09-06
Publication Date
2025-06-25
Estimated Expiration
2037-08-04

AI Technical Summary

Technical Problem

Existing media playback systems lack advanced voice control capabilities for networked devices, limiting the ability to seamlessly integrate and control multiple playback devices and services within a home or commercial environment.

Method used

Implementing networked microphone devices (NMDs) that can receive voice inputs, identify and process commands through various voice services, and synchronize playback across multiple zones, allowing for dynamic configuration and control of media playback systems using voice commands.

Benefits of technology

Enables seamless voice-controlled media playback across multiple zones, enhancing user experience by allowing intuitive control of audio content and device interactions within a networked environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698681000001
    Figure 0007698681000001
  • Figure 0007698681000002
    Figure 0007698681000002
  • Figure 0007698681000003
    Figure 0007698681000003
Patent Text Reader

Abstract

To provide a method, network microphone device, and media playback system for improving listening experience in audio playback between multiple network devices.SOLUTION: A method of causing a voice service to process voice input includes the steps of: receiving voice data indicating voice input via a microphone of a network microphone device; identifying a specific voice service for processing the voice input indicated in the received voice data; and causing the identified voice service to process the voice input via a network interface.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Cross - reference to related applications

[0001] This application claims priority based on U.S. Patent Application No. 15 / 229,868, filed on August 5, 2016, the content of which is hereby incorporated by reference in its entirety into this specification.

Technical Field

[0002] This application relates to consumer products, and in particular, to methods, systems, products, functions, services, and other elements directed to media playback, and to some aspects thereof.

Background Art

[0003] In 2003, Sonos, Inc. filed a patent application titled "Method for Synchronizing Audio Playback Among Multiple Network Devices", which was one of the first patent applications. Until the start of selling media playback systems in 2005, the options for accessing and auditioning digital audio in an out - loud setting were limited. With the Sonos wireless HiFi system, people can experience music from many sources through one or more network playback devices. Through a software control application installed on a smartphone, tablet, or computer, people can play the music they want in any room equipped with a network playback device. Also, for example, using a controller, different songs can be streamed to each room equipped with a playback device, multiple rooms can be grouped for synchronous playback, or the same song can be listened to synchronously in all rooms.

[0004] Considering the increasing interest in digital media so far, there is a need to further release consumer - accessible technologies that can further improve the audition experience.

[0005] The features, aspects, and advantages of the technology disclosed in this specification will be more readily understood with reference to the following description, the appended claims, and the accompanying drawings.

Brief Description of the Drawings

[0006]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Modes for Carrying Out the Invention

[0007] The drawings are intended to illustrate several exemplary embodiments, but it is understood that the present invention is not limited to the arrangements and means shown in the drawings.

[0008] I. Overview

[0009] By using networked microphone devices (NMDs), it is possible to control the home while using voice control. The NMD can be, for example, a Sonos® playback device, server, or system that can receive voice input via a microphone. The NMD can also be a device other than a Sonos® playback device, server, or system that can receive voice input via a microphone (e.g., Amazon®'s Echo®, Apple®'s iPhones®). U.S. Patent Application No. 15 / 098,867, entitled "Designation of Default Playback Device," is incorporated herein by reference, which provides an example of a voice-activated home architecture. Voice control can be beneficial for various devices with "smart" home functions, including playback devices, wireless lighting devices, thermostats, door locks, home automation, and other examples.

[0010] In one embodiment, the voice input detected by the NMD is sent to a voice service for processing. The NMD, such as a playback device, may function as a microphone interface or speaker interface for this voice service. The voice input is detected by the NMD's microphone and then sent to a specific voice service for processing. The voice service can then return a command or other result of the voice input.

[0011] A particular voice service may be selected for the media playback system, possibly during the setup procedure. The user may select the same voice service that they are using on their smartphone or tablet computer (or other personal electronic device), perhaps because they are familiar with that voice service or because they want to use the same controls on the playback device as they do on their smartphone for a similar experience. If a particular voice service is set up on the user's smartphone, the smartphone can send the setup information for that voice service (e.g., user authentication information) to the NMD to facilitate automatic setup of that voice service on the NMD.

[0012] Optionally, multiple voice services may be set up for the NMD, or the NMD's system (e.g., a media playback system with multiple playback devices). One or more services may be set up during the setup procedure. Additional voice services may be set up for the system later. Thus, the NMD described herein may function as an interface for multiple voice services and may potentially reduce the need for an NMD for each voice service to interact with each voice service. Additionally, the NMD may operate in cooperation with service-specific NMDs present in the home to process certain voice commands.

[0013] If two or more voice services are set up for the NMD, a particular voice service can be activated by uttering the activation work corresponding to that particular voice service. For example, when querying Amazon's service, the user may utter the wake word "Alexa" followed by voice input. Other examples include "Okay, Google" when querying Google's service and "Hey, Siri" when querying Apple's service.

[0014] Alternatively, if no wake word is used for a given voice input, the NMD can identify the voice service for processing that voice input. In some cases, the NMD may identify the default voice service. Alternatively, the NMD may identify a specific voice service based on context. For example, the NMD may use the voice service that was most recently queried, based on the assumption that the user wishes to use the same voice service again. Other examples are possible.

[0015] As described above, voice input to the NMD may sometimes be indicated using a general wake word. In some cases, this may be a manufacturer-specific wake word rather than a wake word associated with any particular voice service (e.g., "Hey, Sonos" if the NMD is a Sonos® playback device). Upon receiving such a wake word, the NMD can identify the specific voice service to process the request. For example, if the voice input following the wake word is related to a particular type of command (e.g., playing music), that voice input may be sent to a specific voice service associated with that type of command (e.g., a music streaming service with voice command capabilities).

[0016] The NMD may, in some cases, send the voice input to multiple voice services and, as a result, obtain respective results from the voice services that were queried. The NMD can evaluate these results and respond with the "best" result (e.g., the result that most closely matches the desired action). For example, if the voice input is "Hey, Sonos, play Taylor Swift songs," the first voice service may respond with search results for "Taylor Swift," while the second voice service may respond with identifiers for audio tracks by the artist Taylor Swift. In that case, the NMD can use the identifier for the Taylor Swift audio track from the second voice service to play Taylor Swift songs in accordance with the voice input.

[0017] As described above, the exemplary technology is related to voice services. An exemplary embodiment may include a step in which the NMD receives voice data indicating a voice input via a microphone. The NMD may identify a voice service for processing the voice input from among a plurality of voice services registered in the media playback system, and cause the identified voice service to process the voice input.

[0018] Another exemplary embodiment may include a step in which the NMD receives input data indicating a command to register one or more voice services in the media playback system. The NMD can detect the voice services registered in the NMD. The NMD may cause the voice services registered in the NMD to be registered in the media playback system.

[0019] A third exemplary embodiment may include a step in which the NMD receives voice data indicating a voice input via a microphone. The NMD may determine that a part of the received voice data indicates a general wake word that does not correspond to a specific voice service. The NMD may cause a plurality of voice services to execute processing of the voice input. The NMD may output a result obtained from a predetermined one of the plurality of voice services.

[0020] Each of these exemplary embodiments may be embodied as a method, a device configured to execute the embodiment, a system of devices configured to execute the embodiment, or a non-transitory computer-readable medium including instructions executable by one or more processors to execute the embodiment. It will be understood by those skilled in the art that the present disclosure includes many other embodiments including combinations of the exemplary features described herein. Also, for purposes of illustration, any of the exemplary operations described as being performed by a given device may be performed by any suitable device including the devices described herein. Further, any device may cause any of the other devices described herein to perform any of the operations described herein.

[0021] Some of the examples described herein refer to functions performed by a given actor such as a "user" and / or other entity, but this is for illustrative purposes only. Such actions by such exemplary actors should not be construed as required unless expressly required by the language of the claims themselves.

[0022] II. Examples of Operating Environments FIG. 1 shows an exemplary configuration of a media playback system 100 that may be implemented or implemented in one or more embodiments disclosed herein. As shown, media playback system 100 is associated with an exemplary home environment having a plurality of rooms and spaces, such as a master bedroom, office, dining room, and living room. As shown in the example of FIG. 1, media playback system 100 includes playback devices 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, 124, control devices 126 and 128, and a wired or wireless network router 130.

[0023] Furthermore, descriptions of the different components of the exemplary media playback system 100 and how the different components act to provide a media experience to the user are set forth in the following sections. Although the description herein generally refers to the media playback system 100, the techniques described herein are not limited to use in the home environment shown in FIG. 1. For example, the techniques described herein are beneficial in environments where multi-zone audio is desired, such as commercial environments like restaurants, malls, or airports, vehicles such as sports utility vehicles (SUVs), buses or cars, ships, or boats, airplanes, etc.

[0024] a. Exemplary Zone Player FIG. 2 shows a functional block diagram of an exemplary playback device 200 that constitutes one or more of the playback devices 102-124 of the media playback system 100 of FIG. 1. The playback device 200 may include a processor 202, a software component 204, a memory 206, an audio processing component 208, an audio amplifier 210, a speaker 212, and a network interface 214. The network interface 214 may include a wireless interface 216, a wired interface 218, and a microphone 220. In some cases, the playback device 200 does not include the speaker 212, but may include a speaker interface for connecting the playback device 200 to an external speaker. In another case, the playback device 200 does not include either the speaker 212 or the audio amplifier 210, but may include an audio interface for connecting the playback device 200 to an external audio amplifier or an audio visual receiver.

[0025] In one example, the processor 202 may be a clock-driven computer component configured to process input data based on instructions stored in the memory 206. The memory 206 may be a non-transitory computer-readable recording medium configured to store instructions executable by the processor 202. For example, the memory 206 may be a data storage capable of loading one or more software components 204 executable by the processor 202 to perform a certain function. In one example, the function may include the step of the playback device 200 reading audio data from an audio source or another playback device. In another example, the function may include the step of the playback device 200 transmitting audio data to another device or playback device on the network. In yet another example, the function may include the step of pairing the playback device 200 with one or more playback devices to create a multi-channel audio environment.

[0026] A certain function includes the step of the playback device 200 synchronizing the playback of audio content with one or more other playback devices. While synchronizing the playback, it is preferable that the listener does not notice the delay between the playback of the audio content by the playback device 200 and the playback by one or more other playback devices. U.S. Patent No. 8,234,395 titled "System and Method for Synchronizing Operations between Multiple Independent Clock Digital Data Processing Devices" is incorporated herein by reference, which provides more detailed examples of synchronizing audio playback between playback devices.

[0027] Furthermore, the memory 206 may be configured to store data. The data may be associated with, for example, a playback device 200 such as a playback device 200 that is included as part of one or more zones and / or zone groups, an audio source accessible by the playback device 200, or a playback queue that can be associated with the playback device 200 (or another playback device). The data may be updated periodically and stored as one or more state variables indicating the state of the playback device 200. Also, the memory 206 may include data associated with the states of other devices of the media system and may be shared among the devices as needed so that one or more devices can have near real-time data related to the system. Other embodiments are possible.

[0028] The audio processing component 208 may include one or more digital-to-analog converters (DACs), audio processing components, audio enhancement components, or digital signal processors (DSPs), etc. In certain embodiments, one or more audio processing components 208 may be sub-components of the processor 202. In certain embodiments, audio content may be processed and / or intentionally modified by the audio processing component 208 to generate an audio signal. The generated audio signal is transmitted to the audio amplifier 210, amplified, and reproduced through the speaker 212. In particular, the audio amplifier 210 may include a device configured to amplify the audio signal to a level capable of driving one or more speakers 212. The speaker 212 may comprise a separate transducer (e.g., a "driver") or a complete speaker system including a housing enclosing one or more drivers. Certain drivers provided in the speaker 212 may include, for example, a subwoofer (e.g., for low frequencies), a mid-range driver (e.g., for mid frequencies), and / or a tweeter (for high frequencies). In some cases, each transducer of one or more speakers 212 may be driven by a corresponding individual audio amplifier of the audio amplifier 210. In addition to generating an analog signal for reproduction by the playback device 200, the audio processing component 208 processes the audio content and transmits the audio content for reproduction by one or more other playback devices.

[0029] The audio content processed and / or reproduced by the playback device 200 may be received via an external source, such as an audio line-in input connection (e.g., an auto-detecting 3.5 mm audio line-in connection) or the network interface 214.

[0030] The network interface 214 may be configured to enable a data flow between the playback device 200 and one or more other devices on the data network. In this way, the playback device 200 may be configured to receive audio content via the data network from one or more other playback devices that communicate with the playback device, network devices within a local area network, or an audio content source on a wide area network such as the Internet, for example. In one example, the audio content and other signals transmitted and received by the playback device 200 may be transmitted in the form of digital packets including an Internet Protocol (IP)-based source address and an IP-based destination address. In such a case, the network interface 214 can appropriately receive and process data destined for the playback device 200 by analyzing the digital packet data.

[0031] As shown, network interface 214 may include a wireless interface 216 and a wired interface 218. The wireless interface 216 provides network interface functionality for the playback device 200 and may wirelessly communicate with other devices (e.g., other playback devices, speakers, receivers, network devices, control devices within the data network associated with the playback device 200) based on a communication protocol (e.g., any of the wireless standards (specifications) including wireless standards IEEE802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G mobile communication standards, etc.). The wired interface 218 provides network interface functionality for the playback device 200 and may communicate via a wired connection with other devices based on a communication protocol (e.g., IEEE802.3). The network interface 214 shown in FIG. 2 includes both a wireless interface 216 and a wired interface 218, but in some embodiments, the network interface 214 may include only a wireless interface or only a wired interface.

[0032] The microphone 220 may be configured to detect sound within the environment of the playback device 200. The microphone may be attached, for example, to the outer wall of the housing of the playback device. The microphone may be any type of microphone currently known or later developed, such as a condenser microphone, an electret condenser microphone, or a dynamic microphone. The microphone may be sensitive to a portion of the frequency range of the speaker 220. One or more of the speakers 220 may operate in reverse to the microphone 220. In some aspects, the playback device 200 may not include the microphone 220.

[0033] In one example, the playback device 200 and another playback device may be paired to play two separate audio components of the audio content. For example, the playback device 200 may be configured to play the left-channel audio component, while the other playback device may be configured to play the right-channel audio component. Thereby, the stereo effect of the audio content can be generated or enhanced. The paired playback devices (also referred to as "coupled playback devices") may further play the audio content in synchronization with other playback devices.

[0034] In another example, the playback device 200 may be acoustically integrated with one or more other playback devices to form a single integrated playback device. The integrated playback device can be configured to process and reproduce sound differently than a non-integrated playback device or a paired playback device. This is because the integrated playback device can add speakers for playing the audio content. For example, if the playback device 200 is designed to play audio content in the low-frequency range (e.g., a subwoofer), the playback device 200 may be integrated with a playback device designed to play audio content in the full-frequency range. In this case, when the full-frequency range playback device is integrated with the low-frequency playback device 200, it may be configured to play only the mid- and high-frequency components of the audio content. On the other hand, the low-frequency range playback device 200 plays the low-frequency components of the audio content. Further, the integrated playback device may be paired with a single playback device or even another integrated playback device.

[0035] As an example, currently, Sonos, Inc. sells and offers playback devices including "PLAY:1", "PLAY:3", "PLAY:5", "PLAYBAR", "CONNECT:AMP", "CONNECT", and "SUB". In any other past, current, and / or future playback devices, the playback devices of the embodiments disclosed herein can be additionally or alternatively implemented and used. Further, it is understood that the playback devices are not limited to the specific examples shown in FIG. 2 or the Sonos products provided. For example, the playback devices may include wired or wireless headphones. In another example, the playback devices may include or interact with docking stations for personal mobile media playback devices. In yet another example, the playback devices may be integrated with another device or component, such as a television, lighting fixture, or some other device for use indoors or outdoors.

[0036] b. Exemplary playback zone configurations Returning to the media playback system of FIG. 1, the environment has one or more playback zones, and each playback zone includes one or more playback devices. The media playback system 100 is formed by one or more playback zones, and one or more zones may be added or removed later to be in the exemplary configuration shown in FIG. 1. Each zone may be given a name based on a different room or space, such as an office, bathroom, master bedroom, bedroom, kitchen, dining room, living room, and / or balcony. In some cases, a single playback zone may include multiple rooms or spaces. In other cases, a single room or space may include multiple playback zones.

[0037] As shown in FIG. 1, each of the zones of the balcony, dining room, kitchen, bathroom, office, and bedroom has one playback device, while each of the zones of the living room and the master bedroom has a plurality of playback devices. The living room zone may be configured such that playback devices 104, 106, 108, and 110 play audio content synchronously either as separate playback devices, as one or more combined playback devices, as one or more integrated playback devices, or any combination thereof. Similarly, in the case of the master bedroom, playback devices 122 and 124 may be configured to play audio content synchronously either as separate playback devices, as a combined playback device, or as an integrated playback device.

[0038] In one example, one or more playback zones in the environment of FIG. 1 are each playing different audio content. For example, a user can listen to hip-hop music played by playback device 102 while grilling in the balcony zone. On the other hand, another user can listen to classical music played by playback device 114 while preparing a meal in the kitchen zone. In another example, a playback zone may play the same audio content in synchronization with another playback zone. For example, when a user is in the office zone, playback device 118 in the office zone may play the same music as that being played by playback device 102 in the balcony. In such a case, since playback devices 102 and 118 are playing the rock music in synchronization, the user can enjoy the audio content played out-loud seamlessly (or at least substantially seamlessly) even when moving between different playback zones. Synchronization between playback zones may be performed in a similar manner as the synchronization between playback devices as described in the aforementioned U.S. Patent No. 8,234,395.

[0039] As described above, the zone configuration of the media playback system 100 may be changed dynamically. In certain embodiments, the media playback system 100 supports multiple configurations. For example, when a user physically moves one or more playback devices into or out of a zone, the media playback system 100 may be reconfigured to accommodate the change. For example, if a user physically moves playback device 102 from a balcony zone to an office zone, the office zone may include both playback device 118 and playback device 102. Optionally, via control devices, such as control devices 126 and 128, playback device 102 may be paired, grouped into the office zone, and / or renamed. On the other hand, if one or more playback devices are moved to an area in a home environment that has not yet been set as a playback zone, a new playback zone may be formed in that area.

[0040] Furthermore, different playback zones of the media playback system 100 may be dynamically combined into zone groups or split into separate playback zones. For example, by combining the dining room zone and the kitchen zone 114 into a zone group for a dinner party, playback devices 112 and 114 can play audio content synchronously. On the other hand, if one user wants to watch TV while another user wants to listen to music in the living room space, the living room zone may be divided into a TV zone including playback device 104 and a listening zone including playback devices 106, 108, and 110.

[0041] c. Exemplary Control Devices Figure 3 shows a functional block diagram of an exemplary control device 300 that constitutes one or both of the control devices 126 and 128 of the media playback system 100. As shown, the control device 300 may include a processor 302, a memory 304, a network interface 306, a user interface 308, a microphone 310, and a software component 312. In one example, the control device 300 may be a control device dedicated to the media playback system 100. In another example, the control device 300 may be a network device installed with media playback system controller application software, such as an iPhone (registered trademark), an iPad (registered trademark), or any other smartphone, tablet, or network device (e.g., a network computer such as a PC or Mac (registered trademark)).

[0042] The processor 302 may be configured to perform functions related to enabling user access, control, and configuration of the media playback system 100. The memory 304 may be a data storage capable of hosting one or more software components that are executed by the processor 302 to perform functions. Also, the memory 304 may be configured to store media playback system controller application software and other data associated with the media playback system 100 and the user.

[0043] In one example, network interface 306 may be based on industrial standards (e.g., infrared, wireless, wired standards such as IEEE802.3, wireless standards such as IEEE802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G communication standards, etc.). In network interface 306, control device 300 may provide means for communicating with other devices within media playback system 100. In one example, data and information (e.g., state variables) may be communicated between control device 300 and other devices via network interface 306. For example, the configuration of playback zones and zone groups in media playback system 100 may be received by control device 300 from a playback device or another network device, or may be transmitted by control device 300 to another playback device or network device via network interface 306. In some cases, the other network device may be another control device.

[0044] Playback device control commands such as volume control and audio playback control may be communicated from control device 300 to the playback device via network interface 306. As described above, changes to the configuration of media playback system 100 can be made by the user using control device 300. Configuration changes may include adding one or more playback devices to a zone, removing one or more playback devices from a zone, adding one or more zones to a zone group, removing one or more zones from a zone group, forming a combined player or an integrated player, splitting a combined player or an integrated player into one or more playback devices, etc. Thus, control device 300 may be referred to as a controller, and control device 300 may be a dedicated controller installed with media playback system controller application software, or may be a network device.

[0045] The control device 300 may include a microphone 310. The microphone 310 may be configured to detect sounds within the environment of the control device 300. The microphone 310 may be any type of microphone known currently or developed in the future, such as a condenser microphone, an electret condenser microphone, a dynamic microphone, etc. The microphone may be sensitive to some frequency ranges. Two or more microphones 310 may be provided to obtain position information of a sound source (e.g., voice, audible sound) and / or assist in filtering background noise.

[0046] The user interface 308 of the control device 300 may be configured to enable user access and control of the media playback system 100 by providing a controller interface such as the controller interface 400 shown in FIG. 4. The controller interface 400 includes a playback control region 410, a playback zone region 420, a playback status region 430, a playback queue region 440, and an audio content source region 450. The illustrated user interface 400 is merely an example of a user interface provided by a network device (and / or the control devices 126 and 128 of FIG. 1) such as the control device 300 of FIG. 3 and is accessed by a user to control a media playback system such as the media playback system 100. Alternatively, other user interfaces in various formats, styles, and interactive sequences may be implemented on one or more network devices to provide similar control access to the media playback system.

[0047] The playback control area 410 may include selectable icons (e.g., by using a touch or cursor). These icons cause the playback devices within the selected playback zone or zone group to play or stop, fast forward, rewind, skip to the next, skip to the previous, turn shuffle mode on / off, turn repeat mode on / off, and turn crossfade mode on / off. The playback control area 410 may include another selectable icon. The other selectable icon may change other settings such as equalization settings, playback volume, etc.

[0048] The playback zone area 420 may include a display of the playback zones within the media playback system 100. In some embodiments, a graphical display of the playback zones may be selectable. By moving additional selectable icons, the playback zones within the media playback system can be managed or configured. For example, other management or configuration such as creating combined zones, creating zone groups, splitting zone groups, and renaming zone groups can be performed.

[0049] For example, as shown in the illustration, the "Group" icon may be provided for each of the graphic displays of the playback zones. The "Group" icon within the graphic display of a certain zone may be selectable to present an option to select one or more zones within the media playback system and group them with a certain zone. Once grouped, the playback devices within the zone grouped with a certain zone are configured to play audio content in synchronization with the playback devices within the certain zone. Similarly, the "Group" icon may be provided within the graphic display of the zone group. In this case, the "Group" icon may be selectable to present an option to deselect one or more zones within the zone group in order to remove one or more zones within the zone group from the zone group. Other interactions for grouping and ungrouping zones may also be possible and implementable via a user interface such as the user interface 400. The display of the playback zones within the playback zone area 420 may be dynamically updated when the playback zone or zone group configuration is changed.

[0050] The playback status area 430 may include a graphic display of the currently playing audio content, the previously played audio content, or the audio content scheduled to be played next within the selected playback zone or zone group. The selectable playback zone or playback group may be visually distinguished on the user interface, for example, within the playback zone area 420 and / or the playback status area 430. The graphic display may include the track title, artist name, album name, album year, track length, and other relevant information beneficial to the user when controlling the media playback system via the user interface 400.

[0051] The playback queue area 440 may include a graphic display of audio content within a playback queue associated with a selected playback zone or zone group. In some embodiments, each playback zone or zone group may be associated with a playback queue that includes information corresponding to zero or more audio items to be played by the playback zone or playback group. For example, each audio item within the playback queue may include a uniform resource identifier (URI), a uniform resource locator (URL), or other identifier usable by a playback device within the playback zone or zone group. These may be used to find and / or retrieve audio items from a local audio content source or a network audio content source and play them on a playback device.

[0052] In one example, a playlist may be added to the playback queue. In this case, information corresponding to each audio item within the playlist may be added to the playback queue. In another example, the audio items within the playback queue may be saved as a playlist. In yet another example, when a playback device is playing continuously streaming audio content, such as Internet radio that plays continuously without stopping, rather than audio items that are not played continuously, such as those having a playback time, the playback queue may be empty or may be "unused" but populated. In another embodiment, the playback queue may include Internet radio and / or other streaming audio content items and may be made "unused" when the playback zone or zone group is playing those items. Other examples are possible.

[0053] When a playback zone or zone group is "grouped" or "ungrouped", the playback queue associated with the affected playback zone or zone group may be cleared or reassociated. For example, if a first playback zone containing a first playback queue is grouped with a second playback zone containing a second playback queue, the resulting zone group may have an associated playback queue. The associated playback queue may initially be empty, or (e.g., if the second playback zone is added to the first playback zone) contain the audio items of the first playback queue, or (e.g., if the first playback zone is added to the second playback zone) contain the audio items of the second playback queue, or combine the audio items of both the first playback queue and the second playback queue. Thereafter, if the resulting zone group is ungrouped, the ungrouped first playback zone may be reassociated with the previous first playback queue, or with an empty new playback queue, or with a new playback queue containing the audio items of the playback queue that was associated with the zone group before the zone group was ungrouped. Similarly, the ungrouped second playback zone may be reassociated with the previous second playback queue, or with an empty new playback queue, or with a new playback queue containing the audio items of the playback queue that was associated with the zone group before the zone group was ungrouped. Other examples are possible.

[0054] Returning to the user interface 400 of FIG. 4, the graphic display of the audio content within the playback queue region 440 may include the track title, artist name, track length, and other relevant information associated with the audio content within the playback queue. In one example, the graphic display of the audio content can be selected and moved with additional selectable icons. This enables the management and / or operation of the playback queue and / or the audio content displayed in the playback queue. For example, the displayed audio content may be removed from the playback queue, moved to a different position within the playback queue, selected to play immediately or after the currently playing audio content, or other operations may be performed. The playback queue associated with a playback zone or zone group may be stored in the memory of one or more playback devices within the playback zone or zone group, the memory of a playback device not within the playback zone or zone group, and / or the memory of other designated devices.

[0055] The audio content source region 450 may include a graphic display of selectable audio content sources. From this audio content source, audio content may be retrieved and played by the selected playback zone or zone group. Explanation of the audio content source can refer to the following sections.

[0056] d. Exemplary Audio Content Sources As previously illustrated, one or more playback devices within a zone or zone group may be configured to retrieve the audio content to be played from a plurality of available audio content sources (e.g., based on the corresponding URI or URL of the audio content). In one example, the audio content may be retrieved directly by the playback device from a corresponding audio content source (e.g., a line-in connection). In another example, the audio content may be provided to the playback device on the network via one or more other playback devices or network devices.

[0057] Exemplary audio content sources may include the memory of one or more playback devices within a media playback system. Examples of media playback systems may include, for example, the media playback system 100 of FIG. 1, a local music library on one or more network devices (e.g., a control device, a network-enabled personal computer, or a network-attached storage (NAS), etc.), a streaming audio service that provides audio content via the Internet (e.g., the cloud), or an audio source connected to the media playback system via a line-in input connection of a playback device or network device, or other possible systems.

[0058] In one embodiment, an audio content source may be periodically added to, or removed from, a media playback system such as the media playback system 100 of FIG. 1. In one example, an index of audio items may be performed each time one or more audio content sources are added, removed, or updated. Indexing the audio items may include scanning for identifiable audio items within all folders / directories shared on the network, where the network is accessible by a playback device within the media playback system. Indexing the audio items may also include creating, or updating, an audio content database that includes metadata (e.g., title, artist, album, track length, etc.) and other related information. The other related information may include, for example, a URI or URL for finding each identifiable audio item. Other examples for managing and maintaining the audio content source are possible.

[0059] The foregoing descriptions of the playback device, control device, playback zone configuration, and media content source are merely examples of an operating environment in which the functions and methods described below can be implemented. Other operating environments and configurations not explicitly described herein with respect to the media playback system, playback device, and network device are equally applicable and may be suitable for implementing the functions and methods.

[0060] e. Multiple exemplary network devices FIG. 5 is a diagram showing a plurality of exemplary devices 500 configured to provide an audio playback experience based on voice control. Those skilled in the art will understand that the devices shown in FIG. 5 are for illustrative purposes only, and that variations including different and / or additional devices may be practicable. As shown, the plurality of devices 500 includes computing devices 504, 506, and 508, network microphone devices (NMDs) 512, 514, and 516, playback devices (PBDs) 532, 534, 536, and 538, and a control device (CR) 522.

[0061] Each of the plurality of devices 500 may be a network-enabled device capable of establishing communication with one or more other devices among the plurality of devices according to one or more network protocols such as NFC, Bluetooth®, Ethernet, and IEEE 802.11 via one or more types of networks such as a wide area network (WAN), a local area network (LAN), and a personal area network (PAN).

[0062] As shown, computing devices 504, 506, and 508 may be part of cloud network 502. Cloud network 502 may include additional computing devices. In one example, computing devices 504, 506, and 508 may be different servers, and in another example, two or more of computing devices 504, 506, and 508 may be modules of a single server. Similarly, each of computing devices 504, 506, and 508 may include one or more modules or servers. For ease of illustration herein, each of computing devices 504, 506, and 508 may be configured to perform a specific function within cloud network 502. For example, computing device 508 may be a source of audio content for a music streaming service.

[0063] As shown, computing device 504 may be configured to interface with NMDs 512, 514, and 516 via communication path 542. NMDs 512, 514, and 516 may be components of one or more "smart home" systems. In some cases, NMDs 512, 514, and 516 may be physically arranged throughout a home, similar to the arrangement of the devices shown in FIG. 1. In other cases, two or more of NMDs 512, 514, and 516 may be physically arranged in relatively close proximity to each other. Communication path 542 may comprise one or more types of networks, such as a WAN, LAN, and / or PAN including the Internet, among others.

[0064] In one example, one or more of NMD512, 514, and 516 may be devices configured to primarily perform voice detection. In another example, one or more of NMD512, 514, and 516 may be components of a device having various major utilities. For example, as described above in connection with FIGS. 2 and 3, one or more of NMD512, 514, and 516 may be the microphone(s) 220 of the playback device 200 or the microphone(s) 310 of the network device 300. Also, in some cases, one or more of NMD512, 514, and 516 may be the playback device 200 or the network device 300. In one example, one or more of NMD512, 514, and / or 516 may include a plurality of microphones arranged in a microphone array.

[0065] As shown, the computing device 506 may be configured to interface with CR522 and PBD532, 534, 536, and 538 via the communication path 544. In one example, CR522 may be a network device such as the network device 200 of FIG. 2. Thus, CR522 may be configured to provide the controller interface 400 of FIG. 4. Similarly, PBD532, 534, 536, and 538 may be playback devices such as the playback device 300 of FIG. 3. For this reason, PBD532, 534, 536, and 538 may be physically arranged throughout the home as shown in FIG. 1. For illustrative purposes, PBD536 and 538 may be part of the coupling zone 530, while PBD532 and 534 may be part of their respective zones to which they belong. As described above, PBD532, 534, 536, and 538 may be dynamically coupled, grouped, decoupled, and ungrouped. The communication path 544 may comprise one or more types of networks such as a WAN, LAN, and / or PAN including the Internet and others.

[0066] In one example, similar to NMD512, 514, and 516, CR522 as well as PBD532, 534, 536, and 538 may also be components of one or more "smart home" systems. In some cases, PBD532, 534, 536, and 538 may be placed throughout the same home as NMD512, 514, and 516. Further, as described above, one or more of PBD532, 534, 536, and 538 may be one or more of NMD512, 514, and 516.

[0067] NMD512, 514, and 516 may be part of a local area network, and communication path 542 may include an access point that links the local area network to which NMD512, 514, and 516 belong to computing device 504 via the WAN (the communication path is not shown). Similarly, each of NMD512, 514, and 516 may communicate with each other via such an access point.

[0068] Similarly, CR522 as well as PBD532, 534, 536, and 538 may be part of a local area network and / or a local playback network, as described in the previous section, and communication path 544 may include an access point that links the local area network and / or local playback network to which CR522 as well as PBD532, 534, 536, and 538 belong to computing device 506 via the WAN. Thus, each of CR522 as well as PBD532, 534, 536, and 538 may also communicate with each other via such an access point.

[0069] In one example, a single access point may include communication paths 542 and 544. In one example, each of NMD512, 514, and 516, CR522, and PBD532, 534, 536, and 538 may access cloud network 502 via the same home access point.

[0070] As shown in FIG. 5, each of NMDs 512, 514, and 516, CR 522, and PBDs 532, 534, 536, and 538 may also communicate directly with one or more of the other devices via communication means 546. The communication means 546 described herein may include one or more forms of communication between devices according to one or more network protocols via one or more types of networks, and / or may include communication via one or more other network devices. For example, the communication means 546 may include, by way of example, one or more of Bluetooth™ (IEEE802.15), NFC, Wireless Direct, and / or proprietary wireless among others.

[0071] In one example, CR 522 may communicate with NMD 532 via Bluetooth™ and communicate with PBD 534 via a different local area network. In another example, NMD 514 may communicate with CR 522 via a different local area network and communicate with PBD 536 via Bluetooth. In yet another example, each of PBDs 532, 534, 536, and 538 may communicate with each other via a local playback network according to the spanning tree protocol, while also communicating with CR 522 via a local area network different from the local playback network. Other examples are possible.

[0072] In some cases, the communication means between NMD512, 514, and 516, CR522, and PBD532, 534, 536, and 538 may vary according to the type of communication between devices, the network state, and / or latency requirements. For example, when initially introducing NMD516 into the home together with PBD532, 534, 536, and 538, the communication means 546 may be used. In some cases, NMD516 may transmit identification information corresponding to NMD516 to PBD538 via NFC, and in response, PBD538 may transmit local area network information to NMD516 via NFC (or some other communication format). However, once NMD516 is installed in the home, the communication means between NMD516 and PBD538 may change. For example, NMD516 may communicate with PBD538 continuously via communication path 542, cloud network 502, and communication path 544. In another example, the NMD and PBD may be prevented from ever communicating via local communication means 546. In yet another example, the NMD and PBD may communicate mainly via local communication means 546. Other examples are possible.

[0073] In an exemplary example, NMD512, 514, and 516 may be configured to receive voice input for controlling PBD532, 534, 536, and 538. Available control commands may include control of any of the aforementioned media playback systems, such as playback volume control, playback transport control, music source selection, and grouping, among others. For example, NMD512 may receive voice input for controlling one or more of PBD532, 534, 536, and 538. In response to receiving the voice input, NMD512 may transmit the voice input to computing device 504 for processing via communication path 542. In one example, computing device 504 may convert the voice input into an equivalent text command, parse the text command to identify the command, and then subsequently transmit the text command to computing device 506. In another example, computing device 504 may convert the voice input into an equivalent text command and then subsequently transmit the text command to computing device 506. Computing device 506 may then parse the text command to identify one or more playback commands.

[0074] For example, if the text command is "Play 'Track 1' by 'Artist 1' from 'Streaming Service 1' in 'Zone 1'", the computing device 506 may identify (i) the URL of "Track 1" by "Artist 1" available from "Streaming Service 1", and (ii) at least one playback device within "Zone 1". In this example, the URL of "Track 1" by "Artist 1" from "Streaming Service 1" may be a URL pointing to the computing device 508, and "Zone 1" may be the combined zone 530. Thus, upon identifying the URL and one or both of the PBDs 536 and 538, the computing device 506 may transmit the identified playback URL to one or both of the PBDs 536 and 538 via the communication path 544. One or both of the PBDs 536 and 538 may, in response, retrieve the audio content from the computing device 508 according to the received URL and initiate playback of "Track 1" by "Artist 1" from "Streaming Service 1".

[0075] Those skilled in the art will understand that the above are merely exemplary examples and that other embodiments are also feasible. In some cases, as described above, the operations performed by one or more of the plurality of devices 500 may be alternatively performed, in part or in whole, by one or more other devices among the plurality of devices 500. For example, the conversion from voice input to a text command may be alternatively performed, in part or in whole, by other devices such as the NMD 512, the computing device 506, the PBD 536, and / or the PBD 538. Similarly, the identification of the URL may be alternatively performed, in part or in whole, by another device or a plurality of devices such as the NMD 512, the computing device 504, the PBD 536, and / or the PBD 538.

[0076] f. Exemplary Network Microphone Device FIG. 6 shows a functional block diagram of an exemplary network microphone device 600 that constitutes one or more of NMD512, 514, and 516 of FIG. 5. As shown, network microphone device 600 includes a processor 602, a memory 604, a microphone array 606, a network interface 608, a user interface 610, a software component 612, and a speaker(s) 614. Those skilled in the art will understand that other configurations and arrangements of network microphone devices are also possible. For example, the network microphone device can alternatively exclude the speaker(s) 614 or have a single microphone instead of the microphone array 606.

[0077] Processor 602 may include one or more processors and / or controllers in the form of a general-purpose processor or controller or a dedicated processor or controller. For example, processing unit 602 may include a microprocessor, a microcontroller, an application-specific integrated circuit, and a digital signal processor, among others. Memory 604 may be a data storage capable of hosting one or more software components that are executed by processor 602 to perform functions. Thus, memory 604 may include one or more non-transitory computer-readable recording media such as random access memory, registers, cache, etc. as examples, and one or more non-volatile recording media such as read-only memory, hard disk drive, solid state drive, flash memory, and / or optical storage device, among others.

[0078] The microphone array 606 may be a plurality of microphones configured to detect sound within the environment of the network microphone device 600. The microphone array 606 may include any type of microphone now known or later developed, such as a condenser microphone, an electret condenser microphone, or a dynamic microphone. In one example, the microphone array may be configured to detect audio from one or more directions relative to the network microphone device. The microphone array 606 may be sensitive to some frequency ranges. In one example, a first subset of the microphone array 606 may be sensitive to a first frequency range, while a second subset of the microphone array may be sensitive to a second frequency range. Further, the microphone array 606 may be provided to obtain position information of an audio source (e.g., voice, audible sound) and / or to assist in filtering background noise. In a particular embodiment, the microphone array may be composed of only a single microphone instead of a plurality of microphones.

[0079] The network interface 608 may be configured to facilitate wireless and / or wired communication among various network devices within the cloud network 502, such as CR522, PBD532 - 538, computing devices 504 - 508, etc., in relation to FIG. 5, and other network microphone devices. For this purpose, the network interface 608 can take any form suitable for performing these functions, examples of which include an Ethernet interface, a serial bus interface (e.g., FireWire, USB2.0, etc.), a chipset and antenna configured to facilitate wireless communication, and / or any other interface that provides wired and / or wireless communication. In one example, the network interface 608 may be based on industrial standards (e.g., infrared, wireless, wired standards such as IEEE802.3, wireless standards such as IEEE802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, etc., 4G communication standards, etc.).

[0080] The user interface 610 of the network microphone device 600 may be configured to facilitate user interaction with the network microphone device. In one example, the user interface 608 may include one or more of a physical button, a touch sensor screen(s) and / or surface(s) provided with a graphical interface, etc., to enable the user to directly input into the network microphone device 600. The user interface 610 may further include one or more lighting and speaker(s) 614 to provide visual and / or auditory feedback to the user. In one example, the network microphone device 600 may be further configured to play audio content via the speaker(s) 614.

[0081] Referring to embodiments 700, 800, and 900 shown in FIGS. 7, 8, and 9, which are some exemplary embodiments here, exemplary embodiments of the techniques described herein are presented respectively. For example, these exemplary embodiments can be implemented within an operating environment that includes the media playback system 100 of FIG. 1, one or more of the playback devices 200 of FIG. 2, or one or more of the control devices 300 of FIG. 3, as well as other devices described herein and / or other suitable devices. Further, the operations illustrated as examples of what is performed by the media playback system may be performed by any suitable device, such as a playback device or a control device of the media playback system. Embodiments 700, 800, and 900 may include one or more operations, functions, or actions, as illustrated by one or more of the blocks shown in FIGS. 7, 8, and 9. Although the blocks are illustrated in order, these blocks may be executed simultaneously and / or in an order different from the order described herein. Also, the various blocks may be combined into fewer blocks, divided into additional blocks, and / or removed based on the desired embodiment.

[0082] Furthermore, for the embodiments disclosed herein, the flowchart shows the functions and operations of one executable embodiment of this embodiment. In this regard, each block can represent a module, segment, or portion of program code that includes one or more instructions to be executed by a processor to implement a specific logical function or step in the process. This program code may be stored in any type of computer-readable medium, such as a storage device including, for example, a disk or a hard drive. Examples of computer-readable media include non-transitory computer-readable media such as computer-readable media that stores short-term data, such as register memory, processor cache, and random access memory (RAM). Furthermore, examples of computer-readable media include non-transitory recording media such as secondary or persistent long-term storage, such as read-only memory (ROM), optical disks or magnetic disks, and compact disk read-only memory (CD-ROM). Also, the computer-readable media may be any other volatile or non-volatile storage system. The computer-readable media can be regarded as, for example, a computer-readable recording medium or a tangible storage device. Furthermore, for the embodiments disclosed herein, each block can represent a circuit wired to perform a specific logical function in the process.

[0083] III. Exemplary Systems and Methods for Launching a Voice Service As described above, in one example, a computing device can use a voice service to process voice commands. Embodiment 700 is an exemplary technique for causing a voice service to process voice input.

[0084] a. Receiving Voice Data Indicating a Voice Input In block 702, embodiment 700 includes the step of receiving audio data indicative of an audio input. For example, an NMD such as NMD600 can receive audio data indicative of an audio input via a microphone. As yet another example, any of playback devices 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, and 124 or control devices 126 and 128 of FIG. 1 may be an NMD and may also receive audio data indicative of an audio input. In yet another example, the NMDs include NMDs 512, 514, and 516, PBDs 532, 534, 536, and 538, and CR 522 of FIG. 5.

[0085] The NMD may continuously record ambient noise (i.e., listen for audio input) via one or more microphones. The NMD may store this continuous recording in a ring buffer or circular buffer. In such a buffer, the recording may be overwritten (i.e., discarded) unless it includes an audio input. This buffer may be stored locally and / or remotely via any of the devices or servers described herein. In such a case, the step of receiving audio data indicative of an audio input may include the step of recording audio data including the audio input in the buffer.

[0086] NMD can detect that an audio input has been received by detecting that a portion of the audio data includes a wake word or wake phrase. For example, the audio input may include a wake word followed by an audio command. The wake word can enable NMD to start a time interval or time frame for actively listening for the audio input. The time interval or time frame may expire after a certain amount of time (e.g., one minute after NMD receives the first audio input). U.S. Patent Application No. 15 / 131,776, entitled "Actions Based on User ID," is incorporated herein by reference and further examples are described therein. Some exemplary commercially used wake words include "Hey, Siri" (Apple Inc.), "Okay, Google" (Google Inc.), and "Alexa" (Amazon Inc.). Alternatively, the wake word may be unique (e.g., user-defined).

[0087] Returning to FIG. 1 for illustration, a user can issue a specific audio input while in the master bedroom zone. The playback device 122 (and / or playback device 124) functioning as NMD can listen for the audio input (i.e., recording via a microphone and perhaps recording to a buffer) and detect the user's voice as the audio input. The specific audio input may include a wake word to enable NMD to easily recognize the user's voice as the audio input.

[0088] Exemplary voice commands may include commands that instruct to change either the control of the media playback system or the playback settings. The playback settings may include, for example, playback volume, playback transport control, music source selection, and grouping among others. Other voice commands may include operations for controlling a TV or playback settings, settings of a mobile phone terminal, or adjusting a lighting device, among other device operations. As more household devices become "smart" (e.g., by being equipped with a network interface), it becomes possible to control various household devices using voice commands.

[0089] As an example, an NMD can receive voice data indicating a voice input via a network interface, perhaps from another NMD within the home. In addition to receiving voice data indicating a voice input via a microphone, the NMD may receive the recording (e.g., if two NMDs are both within the detection range of the voice input).

[0090] In such embodiments, the NMD may not continuously record ambient noise. Rather, in some cases, the NMD may receive a voice input or an instruction that instructs the NMD to "activate" and start recording a voice input or command. For example, a first NMD (e.g., the playback device 104 shown in FIG. 1) may receive a voice input and, in certain situations described herein, send an instruction or instructions that instruct to start recording to one or more second NMDs (e.g., playback devices 106 and / or 108 among others).

[0091] In some instances, before the NMD device receives the voice data, voice recordings from multiple NMDs may be vetted, processed, and / or consolidated into a single voice input. As an example, the NMD512 may be able to receive voice recordings from one or more other NMDs such as 514 or 516. In some embodiments, the PBD532, 534, 536, and / or 538 may be configured as NMDs, and the NMD512 may receive a voice recording from one of the PBD532, 534, 536, and / or 538. The NMD (or NMDs) may vet, process, and / or consolidate the voice recordings into a single voice input and transmit this single voice input to a computing device for further processing.

[0092] b. Identification of the voice service(s) for processing the voice input In block 704, embodiment 700 includes the step of identifying one or more voice services for processing the voice input. For example, the NMD can identify a particular voice service for processing the voice input indicated in the received voice data. Alternatively, the NMD may identify multiple voice services for processing the voice input.

[0093] The NMD can identify a particular voice service for processing the voice input from among the available voice services. The voice services may be made available in the NMD using various techniques. The available voice services may include voice services registered in the NMD. The operation of registering a given voice service in the NMD may include the step of providing user authentication information (e.g., username and password) of the voice service to the NMD and / or the step of providing an identifier of the NMD to the voice service. In such a registration operation, the NMD may be configured to receive the voice input instead of the voice service, and perhaps the voice service may be configured to receive the voice input from the NMD for processing purposes. The registration operation may be performed during the setup procedure.

[0094] In some cases, the NMD may be associated with a media playback system. The NMD can function as part of the media playback system itself (e.g., as a control device or a playback device), or as another device interconnected with the media playback system, and in some cases can facilitate specific operations of the media playback system (e.g., voice control of the playback device). One or more voice services may be registered with a given media playback system, and the NMD can identify the registered voice services in order to process voice inputs.

[0095] In the registration operation of the media playback system, the NMD (e.g., a control device, a playback device, or other related device) of the media playback system may be configured to receive voice inputs instead of voice services. Further, in such registration operations, the voice services may be configured to receive voice inputs from these devices for processing purposes. The operation of registering voice services with the media playback system may be performed during the setup procedure. Exemplary setup procedures include procedures for setting up a playback device (or multiple playback devices) and / or a control device in a new media playback system. Other exemplary setup procedures include procedures for changing the media playback system (e.g., procedures for adding or removing devices from the system, or procedures for setting up voice services in the system).

[0096] In some cases, a single voice service may be available in the NMD, which can facilitate the identification of the voice service for processing voice inputs. The voice input received by the NMD may be directly transmitted to the voice service, or a response may be provided by the NMD. In such embodiments, the NMD will function as a microphone interface and a speaker interface for a single voice service.

[0097] In other cases, multiple voice services may be available in the NMD for processing voice input. In such cases, the NMD can identify a specific voice service for processing voice input from among the multiple voice services. For example, the NMD can identify a specific voice service from among the multiple voice services registered in the media playback system. As described above, the NMD may be part of the media playback system (e.g., as a playback device or a control device), or may be associated with this system.

[0098] The step of identifying a specific voice service for processing voice input may be based on a wake word or wake phrase in the voice input. For example, after receiving voice data indicating voice input, the NMD can determine that a part of the voice data represents a specific wake word. Further, the NMD may determine that this specific wake word corresponds to a specific voice service. In other words, the NMD may determine that a specific wake word or wake phrase is used to activate a specific voice service. For example, examples of specific wake words include "Hey, Siri" for activating the voice service of Apple Inc., "Okay, Google" for activating the voice service of Google Inc., "Alexa" for activating the voice service of Amazon Inc., or "Hey, Cortana" for activating the voice service of Microsoft. Alternatively, a unique wake word (e.g., user-defined) can be defined to activate a specific voice service. If the NMD determines that a specific wake word in the received voice data corresponds to a specific voice service, the NMD can identify that specific voice service as the voice service for processing the voice input in the voice data.

[0099] The step of determining that a specific wake word corresponds to a specific voice service may include the step of executing a query on one or more voice services using voice data (e.g., a part of the voice data corresponding to the wake word or wake phrase). For example, the voice service may provide an application programming interface that can be called by the NMD to determine whether the voice data contains a wake word or wake phrase corresponding to that voice service. The NMD can call the API by sending a specific query regarding the voice service together with the data representing the wake word part in the received voice data to the voice service. Alternatively, the NMD can call its own API. By registering the voice service with the NMD or the media playback system, the API of the voice service or other architecture can be integrated with the NMD.

[0100] If multiple voice services are available in the NMD, the NMD may execute a query with the wake word detection algorithm corresponding to each voice service among the multiple voice services. As described above, the step of executing a query with such a detection algorithm may include the step of calling the respective APIs of the multiple voice services locally on the NMD or remotely using a network interface. In response to the query regarding the wake word detection algorithm of a predetermined voice service, the NMD can receive a response indicating whether the voice data in the query contains a wake word corresponding to that voice service. If the wake word detection algorithm of a specific voice service detects that the received voice data represents a specific wake word corresponding to that voice service, the NMD may select that specific voice service as the voice service for processing the voice input.

[0101] In some cases, the received voice data may contain voice input even though it does not include a recognizable wake word corresponding to a specific voice service. Such a situation may occur when a predetermined wake word is not clearly detected due to ambient noise or other factors, and as a result, the wake word detection algorithm(s) may not recognize a predetermined wake word as corresponding to any specific voice service. Alternatively, it is also possible that the user has not uttered a wake word corresponding to a specific voice service. For example, there may be a case where voice input processing is invoked using a general wake word that does not correspond to a specific voice service (e.g., "Hey, Sonos").

[0102] In such a case, the NMD can identify a default voice service for processing the voice input based on the context. The default voice service may be determined in advance (e.g., set during a setup procedure such as the exemplary procedure described above). In that case, when the NMD determines that the received voice data does not include a wake word corresponding to a specific voice service (e.g., when the NMD does not detect a wake word corresponding to a specific voice service in the voice data), the NMD can select the default voice service for processing the voice input.

[0103] As described above, some exemplary systems may include multiple NMDs installed in multiple zones (e.g., the media playback system 100 of FIG. 1 targeting the living room, kitchen, dining room, and bedroom zones, each having its own playback device). In such a system, the default voice service may be set for each NMD or for each zone. In that case, the voice input detected by a given NMD or zone may be processed by the default voice service of that NMD or zone. In some cases, the NMD may assume that the voice input detected by a given NMD or zone is intended to be processed by the voice service associated with that zone. However, in other cases, the voice input may be sent to a specific NMD or zone based on a wake word or wake phrase (e.g., in the case of "Hey, kitchen", the voice input is sent to the kitchen zone).

[0104] Referring to FIG. 1 for illustration, the playback device 122 and / or 124 may function as the NMD for the master bedroom zone. The voice input detected by and / or sent to this zone (e.g., "Hey master bedroom, what's the weather today?") may be processed by the default voice service of the master bedroom zone. For example, if the default voice service for the master bedroom zone is "Alexa (registered trademark) by Amazon (registered trademark)", at least one of the NMDs in the master bedroom zone will execute a weather-related query with Alexa. If the voice input includes a wake word or wake phrase corresponding to a specific voice service, the default voice service is disabled by that wake word or wake phrase (if the specific voice service is different from the default voice service), and the NMD can identify that specific voice service to process the voice input.

[0105] In some embodiments, the NMD may identify a voice service based on the identification information of the user providing the voice input. A human voice can vary by pitch, voice quality, and other characteristics, which may provide characteristics for identifying a particular user by that user's voice. In some cases, the NMD may be trained to recognize the voices of the users in the household.

[0106] Users in the household may each use their own preferred voice service. For example, the first and second users in the household may set the NMD to use the first voice service and the second voice service, respectively (e.g., SIRI (registered trademark) and CORTANA (registered trademark)). When the NMD recognizes the voice of the first user in a voice input, the NMD may identify the first voice service to process the voice command. However, when the NMD recognizes the voice of the second user in a voice input, the NMD can alternatively identify the second voice service to process the voice command.

[0107] Alternatively, based on the context, NMD may identify a specific voice service for processing the voice input. For example, NMD may identify a specific voice service based on the type of command. NMD (e.g., NMD associated with a media playback system) can recognize certain commands (e.g., play, stop, fast forward) as a specific type of command (e.g., media playback command). In such a case, when NMD determines that the voice input contains a specific type of command (e.g., media playback command), it may identify a specific voice service configured to process that type of command as the voice service for processing the voice input. Further exemplified, a search query may be another exemplary type of command (e.g., "What's the weather today?" or "Where is David Bowie's birthplace?"). When NMD determines that the voice input contains a search query, it may identify a specific voice service (e.g., "GOOGLE") to process the voice input containing the search query.

[0108] In some cases, NMD may determine that the voice input contains a voice command targeted at a specific type of device. In such a case, NMD may identify a specific voice service configured to process voice input targeted at that type of device for processing the voice input. For example, NMD determines that a given voice input is targeted at one or more wireless lighting devices (e.g., "Turn on the light here" is targeted at a "smart" light bulb in the same room as NMD), and may identify a specific voice service configured to process voice input targeted at wireless lighting devices as the voice service for processing the voice input. As another example, NMD determines that a given voice input is targeted at a playback device, and may identify a specific voice service configured to process voice input targeted at the playback device as the voice service for processing the voice input.

[0109] In some instances, NMD can identify a particular voice service to process the voice input based on previous inputs. When the first voice input has been processed by a given voice service, the user may expect that a subsequent second voice input, among other possible contextual elements, is also targeted at the same kind of device and, in particular, the same device, or is provided immediately after the first command, and is thus also processed by the voice service. For example, NMD can determine that the previous voice input has been processed by a given voice service and that the current voice input is targeted at the same kind of operation as the previous voice input (e.g., determines that both are media playback commands). In such a situation, NMD may identify the voice service to process the current voice input.

[0110] As another example, NMD can determine that the previous voice input has been processed by a given voice service and that the current voice input has been received within a threshold time after the previous voice input (e.g., within 1 - 2 minutes). By way of illustration, playback device 114 receives a first voice input ("Hey Kitchen, play Janis Joplin") and identifies a voice service to process the first voice input, such that playback device 114 can play an audio track by Janis Joplin. Subsequently, playback device 114 receives a subsequent second voice input ("Turn up the volume") and may identify a voice service to process the second voice input. Given the similarity between this kind of command as a media playback command and / or the elapsed time between the two voice inputs, playback device 114 may identify the same voice service identified to process the first voice input to process the second voice input.

[0111] As an example, the NMD may identify a first voice service to process the voice input, and then determine that the first voice service is unavailable for processing the voice input (perhaps because a result cannot be received within a certain period of time). The voice service may become unavailable for several reasons, including expiration of the service, technical problems related to cloud services, or malicious events that violate availability (e.g., distributed service disruption attacks).

[0112] In such a case, the NMD can identify an alternative second voice service to process the voice input. This alternative voice service may be the default voice service. Alternatively, multiple voice services registered in the system may be ranked by priority, and this alternative voice service may be the next highest priority voice service. Other examples are possible.

[0113] In some cases, the NMD may request input from the user when identifying an alternative voice service. For example, the NMD may request that the user specify an alternative voice service (e.g., "GOOGLE (TM) is currently unavailable. Do you want to search for another service?"). Additionally, the NMD may identify an alternative voice service and ask the user if they want to search for this alternative voice service instead (e.g., "SIRI (TM) is currently unavailable. Do you want to search for ALEXA (TM) instead?"). Or, as another example, the NMD may notify the user when it executes a query against an alternative voice service and returns results (e.g., "CORTANA (TM) was unavailable. The following results were obtained from SIRI (TM)"). Once the original voice service becomes available again, the NMD may notify the user of this change in situation and perhaps change the current voice service (e.g., "SIRI (TM) is currently available. Do you want to query SIRI (TM) instead?"). Such responses may be generated from audio data stored on the NMD's data storage or from audio data accessible to the NMD.

[0114] When executing a query against an alternative second voice service, the NMD may attempt to apply one or more setting values of the first voice service to the second voice service. For example, if the query is to play media content by a specific artist and the default audio service is set for the first voice service (e.g., a specific media streaming service), the NMD may attempt to execute a query against the second voice service for audio tracks by the specific artist from the default audio service. However, if different setting values (e.g., different default services) are set for the second voice service, such setting values may overwrite the setting values of the first voice service when executing a query against the second voice service.

[0115] In some cases, only a single voice service is available in the NMD. For example, during the configuration of the media playback system, a specific voice service may be selected for the media playback system. As an example, when a specific voice service is selected, the wake words corresponding to other voice services become inactive, and even if these wake words are detected, the processing may not be started. The voice service may include various setting values for changing the operation of the voice service when a query is executed with a voice input. For example, a preferred media streaming service or a default media streaming service can be set. A media playback voice command (e.g., "Play a song by Katy Perry") will refer to media content (e.g., an audio track by Katy Perry) from that specific music service.

[0116] c. Execution of voice input processing by the identified voice service(s) In block 706, embodiment 700 includes the step of causing the identified voice service(s) to process the voice input. For example, the NMD may transmit, via a network interface, data indicating the voice input and a command or query instructing one or more servers of the identified voice service(s) to process the data indicating the voice input. This command or query may cause the identified voice service(s) to process a voice command. This command or query may vary according to the identified voice service(s) so that they conform to the identified voice service (e.g., to the API of the voice service).

[0117] As described above, the voice data may indicate a voice input, which may include a first part representing a wake word and a second part representing a voice command. In some cases, the NMD may transmit only the data indicating at least the second part (e.g., the part representing the voice command) in the voice input. By not including the first part, the NMD can reduce, among other possible advantages, the bandwidth required to transmit the command and avoid misprocessing of voice inputs that may occur due to the wake word. Alternatively, the NMD may transmit data indicating both parts in the voice input or some other part of the voice data.

[0118] After causing the identified voice service to process the voice input, the NMD can receive the result of that processing. For example, if the voice input indicated a search query, the NMD may receive search results. As another example, if the voice input indicated a command for a device (e.g., a media play command for a playback device), the NMD may receive the command and perhaps additional data associated with that command (e.g., the source of the media associated with the command). The NMD can appropriately output these results according to the type of command and the received result.

[0119] Alternatively, if the voice command targets a device other than the NMD, the result may be sent to that device instead of the NMD. For example, referring to FIG. 1, the playback device 114 in the kitchen zone may receive a voice input (e.g., for adjusting media playback on the playback device 112) that targets the playback device 112 in the dining room zone. In such an embodiment, the playback device 114 may smoothly process the voice input, but the result of this processing (e.g., a command to adjust media playback may be sent to the playback device 112). Alternatively, the voice service may send the result to the playback device 114, the playback device 114 may send the command to the playback device 112, or the playback device 112 may be made to execute the command.

[0120] The NMD can cause the identified voice service to process some voice inputs, but other voice inputs may be processed by the NMD itself. For example, if the NMD is a playback device, a control device, or other device of a media playback system, the NMD may include voice recognition of media playback commands. As another example, the NMD may process the wake word portion of the voice input. In some cases, processing by the NMD can result in a faster response time than processing using the voice service. However, in some cases, processing using the voice service may result in more effective results and / or results that cannot be obtained by processing via the NMD. In some embodiments, a voice service associated with the NMD (e.g., operated by the manufacturer of the NMD) can easily perform such voice recognition.

[0121] IV. Exemplary Systems and Methods for Activating a Voice Service As described above, in one example, a computing device can use a voice service to process voice commands. Embodiment 800 is an exemplary technique for causing a voice service to process a voice input.

[0122] a. Receiving audio data indicating an audio input In block 802, embodiment 800 includes the step of receiving audio data indicating an audio input. For example, the NMD can receive, among other executable embodiments, audio data indicating an audio input via a microphone using any of the exemplary techniques described above in connection with block 702 of embodiment 700.

[0123] b. Determining if the received audio data contains a portion representing a general wake word In block 804, embodiment 800 includes the step of determining that the received audio data contains a portion representing a general wake word. A general wake word may not be specific to a particular audio service. Instead, a general wake word may generally correspond to the NMD or a media playback system (e.g., "Hey, Sonos" in the case of a Sonos® media playback system, or "Hey, Kitchen" in the case of the kitchen zone of a media playback system). By being general, it is assumed that a particular audio service may not be launched by the general wake word. Rather, if multiple audio services are registered, it may be assumed that the general wake word will launch all of these audio services to obtain the best results. Alternatively, if a single audio service is registered, it may be assumed that the general wake word will launch that audio service.

[0124] c. Executing audio input processing by one or more audio services In block 806, embodiment 800 includes the step of causing one or more audio services to process the audio input. For example, the NMD can cause, among other executable embodiments, one or more audio services to process the audio input using any of the exemplary techniques described above in connection with block 706 of embodiment 700.

[0125] In some cases, multiple voice services are available for use in the NMD. For example, multiple voice services are registered in a media playback system associated with the NMD. In such an example, the NMD may cause each of the available voice services to process the voice input. For example, the NMD may transmit, via a network interface, data indicating the voice input and a command or query instructing each of the servers of the multiple voice services (if multiple) to process the data indicating the voice input. This command or query may cause the identified voice service(s) to process the voice command. This command or query may vary according to each voice service so that they are compatible with the voice service (e.g., the API of the voice service).

[0126] After causing the voice service(s) to process the voice input, the NMD can receive the result of the processing. For example, if the voice input indicates a search query or a media playback command, the NMD may receive the search result or the command, respectively. The NMD may receive the result from each voice service or a subset of the voice services. In some voice services, results may not be returned for all possible inputs that may occur.

[0127] d. Output result from a specific voice service among the voice service(s) (if multiple) In block 806, embodiment 800 includes the step of outputting the result from a specific voice service among the voice service(s) (if multiple). If a result is received from only one voice service, the NMD may output that result. However, if results are received from multiple voice services, the NMD may select a specific result from among the respective results from the multiple voice services and output that result.

[0128] As an example, in one instance, the NMD may receive an audio input such as "Hey Kitchen, play Taylor Swift's song". Since the wake word portion of the audio input ("Hey, Kitchen") does not specify a particular audio service, the NMD may determine that it is general. When receiving this type of wake word, the NMD may have the audio input processed by multiple audio services. However, if the wake word portion of the audio input includes a wake word corresponding to a particular audio service (e.g., "Hey, Siri"), the NMD may instead have the audio input processed by only the corresponding audio service.

[0129] After having the audio input processed by multiple audio services, the NMD can receive the respective results from these multiple audio services. For example, for the audio command "play Taylor Swift's song", the NMD may receive Taylor Swift's audio track from a first audio service (e.g., ALEXA (registered trademark)) and receive search results related to Taylor Swift from a second audio service (e.g., GOOGLE (registered trademark)). Since the command was to "play" Taylor Swift's song, the NMD may select the audio track from the first audio service rather than the search results from the second audio service. The NMD may output this result by starting the playback of the audio track in the kitchen zone.

[0130] In another example, the audio service related to the processing task may be specific to a particular type of command. For example, a media streaming service (e.g., SPOTIFY (registered trademark)) may have an audio service component for commands related to audio playback. In one instance, the NMD may receive an audio input such as "What's the weather?". For this input, the audio service of the media streaming service may not return a useful result (e.g., a null result or an error result). The NMD may have the possibility of selecting a result from another audio service.

[0131] V. Exemplary Systems and Methods for Registering Voice Services As described above, in one example, a computing device can register one or more voice services to process voice commands. Embodiment 900 is an exemplary technique for causing an NMD to register at least one voice service.

[0132] a. Receiving Input Data Indicating a Command to Instruct to Register One or More Voice Services At block 902, embodiment 900 includes receiving input data indicating a command to instruct one or more second devices to register one or more voice services. For example, a first device (e.g., an NMD) may receive, via a user interface (e.g., a touch screen), input data indicating a command to instruct a media playback system including one or more playback devices to register one or more voice services. In one example, the NMD receives the input as part of the procedure for configuring the media playback system using any of the exemplary techniques described above, especially in relation to block 702 of embodiment 700.

[0133] b. Detecting Voice Services Registered with the NMD At block 904, embodiment 900 includes detecting one or more voice services registered with a first device (e.g., an NMD). Such voice services may include voice services installed on the NMD or specific to the NMD (e.g., part of the NMD's operating system).

[0134] For example, if the NMD is a smartphone or a tablet, it may have one or more applications (an "app") installed that interface with a voice service. The NMD can detect these applications using any suitable technique. Such techniques may vary depending on the manufacturer of the NMD or the operating system. In one example, the NMD may compare a list or database of installed applications with a list of supported voice services to determine which of the voice services being installed on the NMD are supported.

[0135] In other examples, the voice service may be specific to the NMD. For example, the voice services of Apple Inc. and Google Inc. may be incorporated into or pre-installed on devices running the iOS and Android operating systems, respectively. Additionally, some distributions customized in these operating systems (e.g., Amazon Inc.'s FireOS) may include a proprietary voice service (e.g., ALEXA).

[0136] c. Performing registration of the detected voice service(s) to the device At block 906, embodiment 900 includes registering at least one of the detected voice services with one or more second devices. For example, the NMD may register at least one of the detected voice services with a media playback system (e.g., media playback system 100 of FIG. 1) including one or more playback devices. The step of registering this voice service may include transmitting, via a network interface, a message indicating authentication information regarding the voice service to the media playback system (i.e., its at least one device). This message may further include a command, request, or other query that instructs the media playback system to register the voice service using the authentication information from the NMD. In this way, one or more of the same voice services registered on the user's NMD (e.g., smartphone) may be registered on the user's media playback system using the same authentication information as the user's NMD, thereby facilitating the registration operation. Other advantages are also possible.

[0137] VI. Conclusion This specification discloses various exemplary systems, methods, apparatuses, and products, among other components, including firmware and / or software executed on hardware. Such examples are to be understood as merely exemplary and should not be regarded as limiting. For example, it is intended that some or all of the aspects or components of such firmware, hardware, and / or software may be implemented solely in hardware, solely in software, solely in firmware, or in any combination of hardware, software, and / or firmware. Accordingly, the examples provided are not the only ways to implement those systems, methods, apparatuses, and / or products.

[0138] (Feature 1) A method comprising: receiving voice data indicating voice input via a microphone; identifying a voice service for processing the voice input from among a plurality of voice services registered in a media playback system; and causing the identified voice service to process the voice input via a network interface.

[0139] (Feature 2) The step of identifying a voice service for processing the voice input includes: determining that a part of the received voice data represents a specific wake word corresponding to a specific voice service; and identifying, as the voice service for processing the voice input, the specific voice service corresponding to the specific wake word, wherein each of the plurality of voice services registered in the media playback system corresponds to a respective wake word. The method according to Feature 1.

[0140] (Feature 3) The step of determining that a part of the received voice data represents a specific wake word corresponding to a specific voice service includes: executing a query using the received voice data against a wake word detection algorithm corresponding to each voice service of the plurality of voice services; and determining that the specific voice service's wake word detection algorithm has detected that a part of the received voice data represents a specific wake word corresponding to the specific voice service. The method according to Feature 2.

[0141] (Feature 4) The step of identifying a voice service for processing the voice input includes: determining that the received voice data does not include any wake word corresponding to a predetermined voice service among the plurality of voice services registered in the media playback system; and based on the determination, identifying a default voice service from among the plurality of voice services as the voice service for processing the voice input. The method according to Feature 1.

[0142] (Feature 5) The step of identifying the voice service for processing the voice input includes: (i) determining that a previous voice input was processed by a specific voice service, and (ii) determining that the voice input was received within a threshold time after the reception of the previous voice input; and based on the determination, identifying the specific voice service that processed the previous voice input as the voice service for processing the voice input. The method according to Feature 1.

[0143] (Feature 6) The step of identifying the voice service for processing the voice input includes: (i) determining that a previous voice input was processed by a specific voice service, and (ii) determining that the voice input is targeted at the same type of operation as the previous voice input; and based on the determination, identifying the specific voice service that processed the previous voice input as the voice service for processing the voice input. The method according to Feature 1.

[0144] (Feature 7) The step of identifying the voice service for processing the voice input includes: determining that the voice input includes a media playback command; and based on the determination, identifying a specific voice service configured to process the media playback command as the voice service for processing the voice input. The method according to Feature 1.

[0145] (Feature 8) The step of identifying the voice service for processing the voice input includes: determining that the voice input is targeted at a wireless lighting device; and based on the determination, identifying a specific voice service configured to process voice inputs targeted at wireless lighting devices as the voice service for processing the voice input. The method according to Feature 1.

[0146] (Feature 9) The step of identifying a voice service for processing the voice input includes: determining that a part of the received voice data represents a general wake word that does not correspond to any specific voice service; and based on the determination, identifying a default voice service from among the plurality of voice services as the voice service for processing the voice input. The method according to Feature 1.

[0147] (Feature 10) The media playback system includes a plurality of zones. The step of identifying a voice service for processing the voice input includes: determining that the voice input is targeted at a specific zone among the plurality of zones; and based on the determination, identifying a specific voice service configured to process the voice input targeted at the specific zone of the media playback system as the voice service for processing the voice input. The method according to Feature 1.

[0148] (Feature 11) The step of identifying a voice service for processing the voice input includes: determining that a part of the received voice data represents a specific wake word corresponding to a first voice service; determining that the first voice service is not currently available when processing the voice input; and identifying a second voice service different from the first voice service as the voice service for processing the voice input. The method according to Feature 1.

[0149] (Feature 12) The voice input includes a first part representing a wake word and a second part representing a voice command. The step of causing the identified voice service to process the voice input includes transmitting, via a network interface, to one or more servers of the identified voice service, a command instructing the processing of (i) data indicating at least the second part in the voice input and (ii) data indicating the voice command. The method according to Feature 1.

[0150] A tangible non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to execute the method according to any one of Features 1 to 12.

[0151] (Feature 14) A device configured to execute the method according to any one of Features 1 to 12.

[0152] (Feature 15) A media playback system configured to execute the method according to any one of Features 1 to 12.

[0153] (Feature 16) The networked microphone device includes: (i) a microphone; (ii) a network interface; (iii) one or more processors; and (iv) a tangible non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, are operable to cause the networked microphone device to execute a method. The method includes: (a) receiving, via the microphone, audio data indicating an audio input; (b) identifying, from a plurality of audio services registered in a media playback system, an audio service for processing the audio input; and (c) causing, via the network interface, the identified audio service to process the audio input.

[0154] (Feature 17) The step of identifying an audio service for processing the audio input includes: (i) determining that a part of the received audio data represents a specific wake word corresponding to a specific audio service; and (ii) identifying the specific audio service corresponding to the specific wake word as the audio service for processing the audio input, where each of the plurality of audio services registered in the media playback system corresponds to a respective wake word. The microphone device according to Feature 16.

[0155] (Feature 18) The step of determining that a part of the received voice data represents a specific wake word corresponding to a specific voice service includes: (i) querying, using the received voice data, a wake word detection algorithm corresponding to each of a plurality of voice services; (ii) determining that the wake word detection algorithm for the specific voice service has detected that a part of the received voice data represents a specific wake word corresponding to the specific voice service. The microphone device according to Feature 17.

[0156] (Feature 19) The step of identifying a voice service for processing voice input includes: (i) determining that the received voice data excludes any wake word corresponding to a predetermined voice service among a plurality of voice services registered in the media playback system; (ii) based on the determination, identifying the default voice service among the plurality of voice services as the voice service for processing the voice input. The microphone device according to Feature 16.

[0157] (Feature 20) The step of identifying a voice service for processing voice input includes: (i) determining (a) that the previous voice input was processed by a specific voice service, and (b) that the next voice input was received within a threshold period after the previous voice input was received; (ii) based on the determination, identifying the specific voice service as the voice service for processing the next voice input. The microphone device according to Feature 16.

[0158] (Feature 21) The step of identifying an audio service for processing an audio input includes: (i) determining (a) that a previous audio input was processed by a specific audio service, and (b) that the next audio input targets the same type of operation as the previous audio input; and (ii) based on this determination, identifying the specific audio service as the audio service for processing the next audio input. The microphone device according to feature 16.

[0159] (Feature 22) The step of identifying an audio service for processing an audio input includes: (i) determining that the audio input includes a media playback command; and (ii) based on this determination, identifying a specific audio service configured to process the media playback command as the audio service for processing the audio input. The microphone device according to feature 16.

[0160] (Feature 23) The step of identifying an audio service for processing an audio input includes: (i) determining that the audio input targets a wireless lighting device; and (ii) based on this determination, identifying a specific audio service configured to process audio inputs targeting wireless lighting devices as the audio service for processing the audio input. The microphone device according to feature 16.

[0161] (Feature 24) The step of identifying an audio service for processing an audio input includes: (i) determining that a part of the received audio data represents a general wake word that does not correspond to any audio service; and (ii) based on this determination, identifying the default audio service among a plurality of audio services as the audio service for processing the audio input. The microphone device according to feature 16.

[0162] (Feature 25) The media playback system includes a plurality of zones. The step of identifying a voice service for processing voice input includes: (i) determining that the voice input is targeted at a specific zone among the plurality of zones; (ii) based on this determination, identifying a specific voice service configured to process the voice input targeted at the specific zone as the voice service for processing the voice input. The microphone device according to Feature 16 includes these steps.

[0163] (Feature 26) The step of identifying a voice service for processing voice input includes: (i) determining that the received voice data represents a specific wake word corresponding to a first voice service; (ii) determining that the first voice service is not currently available for processing the voice input; (iii) identifying a second voice service different from the first voice service as the voice service for processing the voice input. The microphone device according to Feature 16 includes these steps.

[0164] (Feature 27) The voice input includes a first part representing a wake word and a second part representing a voice command. The step of causing the identified voice service to process the voice input includes transmitting, via a network interface to one or more servers of the identified voice service: (i) data representing at least the second part of the voice input; and (ii) a command instructing the processing of the data. The microphone device according to Feature 16 includes these steps.

[0165] (Feature 28) A tangible non-transitory computer-readable medium stores instructions that, when executed by one or more processors, are operative to cause a method in a networked microphone device to be executed, the method including: (i) receiving, via a microphone, audio data indicative of an audio input; (ii) identifying, from a plurality of audio services registered in a media playback system, an audio service for processing the audio input; and (iii) causing, via a network interface, the identified audio service to process the audio input.

[0166] (Feature 29) The step of identifying an audio service for processing the audio input includes: (i) determining that a portion of the received audio data represents a specific wake word corresponding to a specific audio service; and (ii) identifying, as the audio service for processing the audio input, the specific audio service corresponding to the specific wake word, wherein each of the plurality of audio services registered in the media playback system corresponds to a respective wake word. The tangible non-transitory computer-readable medium according to Feature 28.

[0167] (Feature 30) The step of determining that a portion of the received audio data represents a specific wake word corresponding to a specific audio service includes: (i) querying, using the received audio data, a wake word detection algorithm corresponding to each of the plurality of audio services; and (ii) determining that the wake word detection algorithm for the specific audio service has detected that a portion of the received audio data represents a specific wake word corresponding to the specific audio service. The tangible non-transitory computer-readable medium according to Feature 29.

[0168] (Feature 31) The step of identifying a voice service for processing a voice input includes: (i) determining that the received voice data excludes any wake word corresponding to a predetermined voice service among a plurality of voice services registered in the media playback system; (ii) based on this determination, identifying the default voice service among the plurality of voice services as the voice service for processing the voice input, in the tangible non-transitory computer-readable medium according to Feature 28.

[0169] (Feature 32) The step of identifying a voice service for processing a voice input includes: (i) determining (a) that a previous voice input was processed by a specific voice service, and (b) that a next voice input was received within a threshold period after the previous voice input was received; (ii) based on this determination, identifying the specific voice service as the voice service for processing the next voice input, in the tangible non-transitory computer-readable medium according to Feature 28.

[0170] (Feature 33) The step of identifying a voice service for processing a voice input includes: (i) determining (a) that a previous voice input was processed by a specific voice service, and (b) that the next voice input targets the same type of operation as the previous voice input; (ii) based on this determination, identifying the specific voice service as the voice service for processing the next voice input, in the tangible non-transitory computer-readable medium according to Feature 28.

[0171] (Feature 34) The step of identifying a voice service for processing a voice input includes: (i) determining that the voice input includes a media playback command; (ii) based on this determination, identifying a specific voice service configured to process the media playback command as the voice service for processing the voice input, in the tangible non-transitory computer-readable medium according to Feature 28.

[0172] (Feature 35) The step of identifying a voice service for processing voice input includes: (i) determining that a part of the received voice data represents a general wake word that does not correspond to any voice service; (ii) based on this determination, identifying the default voice service among a plurality of voice services as the voice service for processing the voice input. The tangible non-transitory computer-readable medium according to feature 28.

[0173] (Feature 36) The media playback system includes a plurality of zones. The step of identifying a voice service for processing voice input includes: (i) determining that the voice input is targeted at a specific zone among the plurality of zones; (ii) based on this determination, identifying a specific voice service configured to process the voice input targeted at the specific zone as the voice service for processing the voice input. The tangible non-transitory computer-readable medium according to feature 28.

[0174] (Feature 37) The step of identifying a voice service for processing voice input includes: (i) determining that the received voice data represents a specific wake word corresponding to a first voice service; (ii) determining that the first voice service is not currently available for processing the voice input; (iii) identifying a second voice service different from the first voice service as the voice service for processing the voice input. The tangible non-transitory computer-readable medium according to feature 28.

[0175] (Feature 38) The voice input includes a first part representing a wake word and a second part representing a voice command. The step of causing the identified voice service to process the voice input includes transmitting, via a network interface, to one or more servers of the identified voice service: (i) data representing at least the second part of the voice input; and (ii) a command instructing the processing of the data. The tangible non-transitory computer-readable medium according to feature 28.

[0176] (Feature 39) (i) Receiving, via a microphone of a networked microphone device, audio data indicative of an audio input; (ii) determining that a portion of the received audio data represents a specific wake word corresponding to a specific audio service among a plurality of audio services registered in the media playback system, wherein each of the plurality of audio services registered in the media playback system corresponds to a respective wake word; (iii) causing the specific audio service to process the audio input via a network interface of the networked microphone device, wherein causing the specific audio service to process the audio input includes transmitting, via the network interface of the microphone device, data indicative of the audio input to one or more servers of the specific audio service. A method.

[0177] Furthermore, references to "embodiments" in this specification mean that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one exemplary embodiment of the invention. The use of this phrase in various parts of the specification does not necessarily refer to the same embodiment, nor are they separate or alternative embodiments mutually exclusive of other embodiments. Thus, it is expressly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0178] This specification is broadly presented with respect to exemplary environments, systems, procedures, steps, logical blocks, processes, and other symbolic representations that are similar in operation to data processing devices directly or indirectly connected to a network. These process descriptions and representations are generally used by those of ordinary skill in the art and can most efficiently convey the substance of their work to other such persons. Numerous specific details are provided in order to understand the present disclosure. However, it will be understood by those of ordinary skill in the art that the specific embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits are not described in detail so as not to unnecessarily obscure the embodiments. Accordingly, the scope of the present disclosure is defined rather by the appended claims than by the above-described embodiments.

[0179] If any of the appended claims is read to cover only an implementation in software and / or firmware, it is expressly defined herein that one or more of the elements in at least one instance include a tangible non-transitory storage medium that stores software and / or firmware, such as a memory, a DVD, a CD, a Blu-ray (registered trademark), etc.

Claims

1. One or more amplifiers configured to drive one or more speakers, At least one microphone array, A network interface, One or more processors, A tangible non - volatile computer - readable medium storing instructions that, when executed by one or more processors, cause a playback device to perform the following steps, Comprising, Continuously capturing audio in one or more buffers via at least one microphone array, Analyzing the captured audio using a plurality of wake - word detection algorithms that are executed simultaneously on one or more processors, where each wake - word detection algorithm corresponds to a respective voice assistant service of a plurality of voice assistant services supported by the playback device, When a particular wake - word detection algorithm among the plurality of wake - word detection algorithms detects a wake - word corresponding to a particular voice assistant service in the captured audio, transmitting the captured audio, via the network interface, to the particular voice assistant service, where the captured audio includes a voice input, and the voice input includes a command to change at least one playback setting of a media playback system including the playback device, After transmitting the captured audio, receiving, via the network interface, from one or more servers of the particular voice assistant service, a command to change at least one playback setting according to the command, Changing at least one playback setting based on the command, Playing at least one audio track via one or more amplifiers configured to drive one or more speakers in a state where at least one playback setting has been changed, A playback device that performs the above.

2. Transmitting a search query to a particular voice assistant service via the network interface, Receiving, in response to a search query, data representing search results from one or more servers of a particular voice assistant service via a network interface, where the search results include an audio track corresponding to the search query, the search results are specific to a particular voice assistant service among a plurality of voice assistant services, and the search results include at least one audio track. The playback device according to claim 1, which performs the above.

3. The captured audio is the first captured audio, and the playback device Further performs the following steps before capturing the first audio. Continuously capturing a second audio in one or more buffers via a microphone array. Analyzing the captured second audio using a plurality of wake word detection algorithms executed simultaneously on one or more processors. Detecting a wake word corresponding to a particular voice assistant service in the captured second audio, where the captured second audio includes a voice command and the voice command includes a search query. The playback device according to claim 2.

4. The playback device is a first playback device. The step of changing at least one playback setting based on a command includes the step of participating in a synchronization group including a second playback device. Further, performing the step of receiving at least one audio track from the second playback device via a network interface. The playback device according to any one of claims 1 to 3.

5. The playback device is a first playback device. The step of changing at least one playback setting based on a command includes the step of forming a synchronization group including the first playback device and the second playback device. The step of playing at least one audio track includes the step of playing at least one audio track in synchronization with the second playback device of the synchronization group. The playback device according to any one of claims 1 to 3.

6. Furthermore, the playback device according to claim 5 performs a step of transmitting at least one audio track to a second playback device via a network interface.

7. The step of changing at least one playback setting based on a command includes a step of selecting a music source of at least one audio track, and the playback device according to any one of claims 1 to 6.

8. Continuously capturing audio in one or more buffers via a microphone array of a playback device, Analyzing the captured audio using a plurality of wake word detection algorithms that are simultaneously executed on one or more processors of the playback device, where each wake word detection algorithm corresponds to each voice assistant service among the plurality of voice assistant services supported by the playback device. When a specific wake word detection algorithm among the plurality of wake word detection algorithms detects a wake word corresponding to a specific voice assistant service in the captured audio, transmitting the captured audio to the specific voice assistant service via the network interface of the playback device, where the captured audio includes a voice input, and the voice input includes a command for changing at least one playback setting of a media playback system including the playback device. After transmitting the captured audio, receiving, via the network interface, a command for changing at least one playback setting according to the command from one or more servers of a specific voice assistant service. Changing at least one playback setting based on the command. Playing at least one audio track via one or more amplifiers configured to drive one or more speakers with at least one playback setting changed. A method comprising.

9. Transmitting a search query to a specific voice assistant service via a network interface. Receiving, in response to a search query, data representing search results from one or more servers of a particular voice assistant service via a network interface, where the search results include an audio track corresponding to the search query, further comprising, The search results are specific to a particular voice assistant service among a plurality of voice assistant services, and the search results include at least one audio track. The method according to claim 8.

10. The captured audio is the first captured audio, Before capturing the first audio, Continuously capturing a second audio in one or more buffers via a microphone array, Analyzing the captured second audio using a plurality of wake word detection algorithms executed simultaneously on one or more processors, Detecting, in the captured second audio, a wake word corresponding to a particular voice assistant service, where the captured second audio includes a voice command, and the voice command includes a search query, further comprising. The method according to claim 9.

11. The playback device is a first playback device, The step of changing at least one playback setting based on a command includes the step of participating in a synchronization group including a second playback device, further comprising receiving at least one audio track from the second playback device via a network interface, The method according to any one of claims 8 to 10.

12. The playback device is a first playback device, The step of changing at least one playback setting based on a command includes the step of forming a synchronization group including the first playback device and the second playback device, The step of playing at least one audio track includes the step of playing at least one audio track in synchronization with the second playback device of the synchronization group, The method according to any one of claims 8 to 10.

13. further comprising transmitting at least one audio track to a second playback device via a network interface. The method according to claim 12.

14. The method according to any one of claims 8 to 13, wherein the step of changing at least one playback setting based on a command includes the step of selecting a music source of at least one audio track.

15. A tangible non - volatile computer - readable medium storing instructions executable by one or more processors such that a playback device executes the method according to any one of claims 8 to 14.

Citation Information

Patent Citations

  • Electronic apparatus, voice recognition system, and voice recognition program

    JP2016024652A

  • Voice recognition client device and server-type voice recognition device

    JP2016095383A

  • Satellite volume control

    US20140363022A1

  • Detecting Self-Generated Wake Expressions

    US20150006176A1