Media playback system with voice assistance
By selectively using a first VAS to process voice inputs in a media playback system, the system addresses the limitations of traditional VAS in handling complex commands and multizone playback, resulting in improved voice control and user experience.
Patent Information
- Application Number
- JP2023144379
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-09-29
- Filing Date
- 2023-09-06
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2038-09-28
AI Technical Summary
Traditional Voice Assistant Services (VAS) face challenges in efficiently processing voice inputs for smart home devices, particularly in handling complex voice commands and multizone media playback systems, due to heavy computational loads and limitations in natural language understanding.
The media playback system selects a first VAS over a second VAS to process voice inputs, using a network microphone device to detect voice commands and determine if they satisfy specific command criteria, thereby activating the first VAS for more advanced voice control functions.
This approach enhances voice control capabilities by accurately processing complex voice commands and supporting advanced features like multizone playback and device grouping, improving the overall user experience in smart home environments.
Smart Images

Figure 0007676491000001 
Figure 0007676491000002 
Figure 0007676491000003
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Patent Application No. 15 / 721,141, filed September 29, 2017, the contents of which are incorporated herein by reference in their entirety. [Technical field]
[0002] This application relates to consumer products, and in particular to methods, systems, products, features, services, and other elements directed to voice control of media playback, or aspects thereof. [Background technology]
[0003] Until Sonos Inc. filed one of its first patent applications in 2003, entitled "Method of Synchronizing Audio Playback Among Multiple Networked Devices," and began selling its media playback system in 2005, options for accessing and listening to digital audio in out-loud settings were limited. Sonos wireless HiFi systems allow people to experience virtually unlimited music from many sources through one or more network playback devices. Through a software control application installed on a smartphone, tablet, or computer, people can play the music they want in every room equipped with a network playback device. In addition, the controller can be used, for example, to stream different songs to each room equipped with a playback device, group multiple rooms for synchronized playback, or listen to the same song in every room in sync. Summary of the Invention
[0004] Given the continuing growth in interest in digital media, there continues to be a need for further development of consumer-accessible technologies that can further enhance the listening experience. [Brief description of the drawings]
[0005] The features, aspects, and advantages of the technology disclosed herein will become better understood with reference to the following description, the appended claims, and the accompanying drawings.
[0006] [Figure 1] FIG. 1 illustrates a media playback system in which some embodiments may be implemented. [Figure 2A] FIG. 2A is a functional block diagram of an exemplary playback device. [Figure 2B] FIG. 2B is an isometric view of an exemplary playback device including a network microphone device. [Figure 3A] FIG. 3A is a diagram illustrating example zones and zone groups according to an embodiment of the present disclosure. [Figure 3B] FIG. 3B is a diagram illustrating example zones and zone groups according to an embodiment of the present disclosure. [Figure 3C] FIG. 3C is a diagram illustrating example zones and zone groups according to an embodiment of the present disclosure. [Figure 3D] FIG. 3D is a diagram illustrating example zones and zone groups according to an embodiment of the present disclosure. [Figure 3E] FIG. 3E is a diagram illustrating example zones and zone groups according to an aspect of the present disclosure. [Figure 4] FIG. 4 is a functional block diagram of an example controller device according to an aspect of the present disclosure. [Figure 4A] FIG. 4A illustrates a controller interface according to an embodiment of the present disclosure. [Figure 4B] FIG. 4B illustrates a controller interface according to an embodiment of the present disclosure. [Figure 5A] FIG. 5A is a functional block diagram of an example network microphone device according to an embodiment of the present disclosure. [Figure 5B] FIG. 5B is a diagram of an example voice input according to an aspect of the present disclosure. [Figure 6]FIG. 6 is a functional block diagram of an exemplary remote computer according to an embodiment of the present disclosure. [Figure 7A] FIG. 7A is a schematic diagram of an exemplary network system according to an embodiment of the present disclosure. [Figure 7B] FIG. 7B illustrates an example message flow implemented by the example network system of FIG. 7A in accordance with an aspect of the present disclosure. [Figure 8A] FIG. 8A is a flow diagram of an example method for invoking a voice assistant service according to an aspect of the present disclosure. [Figure 8B] FIG. 8B is a block diagram of an example set of command information according to an aspect of the present disclosure. [Figure 9A] FIG. 9A is a table of example voice input commands and associated information according to an aspect of the present disclosure. [Figure 9B] FIG. 9B is a table of example voice input commands and associated information according to an aspect of the present disclosure. [Figure 9C] FIG. 9C is a table of example voice input commands and associated information according to an aspect of the present disclosure. [Figure 10A] FIG. 10A illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 10B] FIG. 10B illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 10C] FIG. 10C illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 11A] FIG. 11A illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 11B] FIG. 11B illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 12A] FIG. 12A illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 12B]FIG. 12B illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 13A] FIG. 13A illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 13B] FIG. 13B illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 14A] FIG. 14A illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 14B] FIG. 14B illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 15A] FIG. 15A illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 15B] FIG. 15B illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 16A] FIG. 16A illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 16B] FIG. 16B illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 17A] FIG. 17A illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 17B] FIG. 17B illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 18A] FIG. 18A illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 18B] FIG. 18B illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 19A] FIG. 19A illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 19B] FIG. 19B illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 20A] FIG. 20A illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure. [Figure 20B] FIG. 20B illustrates an example voice input for invoking a VAS according to an embodiment of the present disclosure.
[0007] The drawings are intended to illustrate some exemplary embodiments, it being understood, however, that the invention is not limited to the arrangement and instrumentality shown in the drawings. To facilitate the description of any particular element, the value of the significant digits in the part numbers refers to the figure in which that element is first introduced. For example, element 107 was first introduced and described with reference to FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] I. Overview Voice control can be beneficial for a "smart home" having smart appliances and associated devices such as wireless lighting devices, home automation devices (e.g., thermostats, door locks, etc.), and audio playback devices. In some embodiments, a network microphone device can be used to control the devices in the smart home. A network microphone device typically includes a microphone for receiving voice input. The network microphone device can forward the voice input to a voice assistant service (VAS). A conventional VAS can be a remote service implemented by a cloud server processing the voice input. The VAS can process the voice input to determine the intent of the voice input. Based on the response, the network microphone device can cause one or more smart devices to take an action. For example, based on the response to an instruction from the VAS, the network microphone device can turn a lighting device on or off.
[0009] The voice input detected by the network microphone device typically includes a wake word followed by an utterance containing the user's request. The wake word is typically a predefined word or phrase used to activate the VAS and trigger it to interpret the intent of the voice input. For example, a user may utter the wake word "Alexa" when querying an AMAZON® VAS. Other examples include "Ok, Google" to trigger a GOOGLE® VAS, and "Hey, Siri" to trigger an APPLE®, or "Hey, Sonos" for a VAS provided by SONOS®.
[0010] The network microphone device listens for a user request or command that accompanies the wake word in the voice input. In some examples, the user request may include a command to control a third-party device, such as a thermostat (e.g., a NEST® thermostat), a lighting device (e.g., a PHILIPS HUE® lighting device), or a media playback device (e.g., a Sonos® playback device). For example, a user may say the wake word "Alexa" followed by "Set the thermostat to 68 degrees" to set the temperature in their home using the AMAZON® VAS. A user may say the same wake word followed by "Turn on the living room" to turn on a lighting device in the living room area of the home. A user may similarly say the wake word followed by a request to play a particular song, album, or playlist of music on a playback device in the home.
[0011] A VAS may employ natural language understanding (NLU) systems to process voice input. NLU systems typically require multiple remote servers programmed to detect the underlying intent of a given voice input. For example, the servers may manage a language lexicon, a parser, grammar and semantic rules, and associated processing algorithms to determine the user's intent.
[0012] One of the challenges faced by traditional VAS is the computationally intensive nature of NLU processing. For example, voice processing algorithms need to be constantly updated to handle nuances in parlance, sentence structure, pronunciation, and other conversational features. Therefore, VAS providers must manage and continuously develop processing algorithms and deploy increasing resources, such as additional cloud servers, to handle the myriad voice inputs received from users around the world.
[0013] One related challenge is that voice control of certain smart devices may require relatively complex voice processing algorithms, placing additional strain on VAS resources. For example, to switch on a set of lighting devices in the living room, one user may prefer to say "flip on the lights," while another user may prefer to say "turn on the living room." Both users have the same intent to turn on the lighting devices, but the structure of the phrase containing the verb is different, not to mention that the latter phrase identifies the living room device while the former does not. To address these challenges, the VAS must dedicate additional resources to deciphering the user's intent, especially when controlling smart devices, which require complex voice processing resources and algorithms, such as algorithms to distinguish subtle yet meaningful variations in the command structure and associated syntax.
[0014] As consumer demand for smart devices increases and these devices become more diverse, some VAS providers may find it difficult to keep up with the advancements. A VAS may have limited system resources and may not be able to respond correctly to incoming voice input. For example, in the above example, the VAS may have the ability to process the utterance "turn on the lights" but may lack the ability to process the utterance "flip on the lights" because the system may be using an algorithm that cannot recognize the intent of the latter, more conventional phrase. In such a case, the user may need to rephrase the original request with more specific information, such as saying "turn on the lights in the living room." Alternatively, the VAS may inform the user that it cannot process such a request, or the VAS may simply ignore the request altogether. In either of these cases, the user may be frustrated by a poor voice control experience.
[0015] For media playback systems, such as multi-zone playback systems, a conventional VAS may be particularly limited. For example, a conventional VAS may only support voice control of basic playback, or may require the user to use certain exaggerated phrasing rather than natural dialogue to interact with the device. Furthermore, a conventional VAS may not support multi-zone playback or other features that a user may want to control, such as device grouping, multi-room volumes, equalizer parameters, and / or audio content for a given playback scenario. Controlling such features may require significantly more resources than are needed for basic playback.
[0016] The media playback system described herein can address these and other limitations of conventional VAS. For example, in some embodiments, the media playback is configured to select a first VAS (e.g., an enhanced VAS) over a second VAS (e.g., a conventional VAS) to process a voice input. In such cases, the media playback system can intervene by selecting the first VAS over the second to process a particular voice input, such as a voice input for controlling other relatively advanced functions of the media playback system. In one aspect, the first VAS can provide improved voice control compared to the voice control provided by the second VAS alone. In some embodiments, at least some voice inputs targeted to the media playback system can be invokable via the second VAS. In these and other embodiments, while at least some voice inputs can be invokable via the second VAS, it may be preferable for the first VAS to process certain voice inputs. For example, the first VAS may be able to process certain voice inputs more reliably and accurately than the second VAS. In some embodiments, the second VAS can be a default VAS to which certain types of voice inputs are typically routed. For example, in some embodiments, a traditional VAS may be better suited to handle requests that include basic Internet queries, such as a voice input of "What's the weather today?" In a related embodiment, a user may use the same wake word (such as "Hey Samantha") to activate both the first and second VAS. In one aspect, when uttering the voice input, the user may be unaware that the selection of the VAS is occurring behind the scenes. In one embodiment, the wake word may be the wake word associated with a traditional VAS, such as AMAZON's ALEXA®.
[0017] In one embodiment, the media playback system may include a network microphone device for capturing voice input, and is configured to (i) capture the voice input via the at least one microphone device, (ii) detect inclusion of one or more commands in the captured voice input, (iii) determine that the one or more commands satisfy corresponding command criteria of the set of command information, and (iv) in response to the determination, (a) select the first VAS over the selection of the second VAS, (b) send the voice input to the first VAS, and (c) process a response to the voice input from the first VAS after sending the voice input.
[0018] In some embodiments, the network microphone device is configured to store the set of command information in a local memory of the network microphone device. In some embodiments, the set of command information may be stored in another network device, such as another network microphone device or a playback device on a local area network (LAN). In some embodiments, the set of command information may be stored on the LAN and / or remotely across multiple network devices. In various embodiments described below, the set of command information may be used in determining whether the media playback system should abandon the selection of the second VAS in favor of the first VAS.
[0019] In some embodiments, the network microphone device may store a list of predefined commands and command criteria associated with the commands. The commands may include, for example, play, control, and zone targeting commands. The command criteria may include, for example, predefined keywords associated with a particular command. A combination of keywords in a voice input may include, for example, speaking the name of a first room in a house (e.g., living room) and speaking the name of a second room in a house (e.g., bedroom). When a user speaks a voice input containing a particular command (e.g., a command to play music) in combination with the keywords, the media playback system selects and activates a first VAS to process the voice input.
[0020] In some embodiments, keywords may be developed through training and adaptive learning algorithms. In certain embodiments, such keywords may be determined on the fly while processing a voice input that includes the keyword. In such cases, the keyword is not predetermined prior to processing the voice input, but allows the first VAS to trigger based on the command. In related embodiments, the keyword may be associated with a particular cognate of a command having the same intent.
[0021] In some embodiments, invoking the first VAS may include sending voice input to a remote server of one or more of the first VAS. In the above example, the first VAS may determine the user's intent to play music in the first and second rooms and respond by causing the media playback system to play the desired audio in the first and second rooms. The first VAS may also instruct the media playback system to form a group consisting of the first and second rooms.
[0022] Although some embodiments described herein may refer to functions being performed by certain actors, such as a "user" and / or other entity, it should be understood that this description is for illustrative purposes only, and no action by any such actor should be construed as required unless the claim itself expressly states so.
[0023] II. Exemplary Operating Environment 1 illustrates an exemplary configuration of a media playback system 100 in which one or more embodiments disclosed herein can be implemented. As illustrated, the media playback system 100 is associated with an exemplary home environment having multiple rooms and spaces, e.g., an office, a dining room, and a living room. Within these rooms and spaces, the media playback system 100 includes playback devices 102 (individually identified as playback devices 102a-102m), network microphone devices 193 (individually identified as "NMDs" 103a-103g), and controller devices 104a, 104b (collectively "controller devices 104"). The home environment may also include other networked devices, such as one or more smart lighting devices 108 and a smart thermostat 110.
[0024] The various playback devices, network microphone devices, and controller devices 102-104 of the media playback system 100, and / or other network devices may be coupled to each other via point-to-point connections and / or other connections, wired and / or wireless, via a LAN including the network router 106. For example, playback device 102j (labeled "left") may have a point-to-point connection with playback device 102a (labeled "right"). In one embodiment, the "left" playback device 102j may communicate with the "right" playback device 102a via a point-to-point connection. In a related embodiment, the "left" playback device 102j may communicate with other network devices via point-to-point connections and / or other connections, such as via a LAN.
[0025] The network router 106 may be coupled to one or more remote computers 105 via a wide area network (WAN). In some embodiments, the remote computer may be a cloud server. The remote computer 105 may be configured to interact with the media playback system 100 in various ways. For example, the remote computer may be configured to facilitate streaming and control of media content playback, such as audio, in a home environment. In one aspect of the technology, which is described in more detail below, the remote computer 105 is configured to provide a first VAS 160 to the media playback system 100.
[0026] In some embodiments, one or more of the playback devices 102 may include an on-board (e.g., integrated) network microphone device. For example, playback devices 102a-e each include a corresponding NMD 103a-e. A playback device that includes a network microphone device may be referred to interchangeably herein as a playback device or a network microphone device, unless otherwise noted in the specification.
[0027] In some embodiments, one or more of the NMDs 103 may be independent devices. For example, NMDs 103f and 103g may be independent network microphone devices. An independent network microphone device may omit components such as a speaker or associated electronics that are typically included in a playback device. In such cases, the independent network microphone device may not generate audio output or may generate limited audio output (e.g., a relatively low quality audio output).
[0028] During use, a network microphone device may receive and process voice input from a nearby user. For example, a network microphone device may capture voice input when it detects that a user has spoken the input. In the illustrated example, NMD 103a of living room playback device 102a may capture voice input from a nearby user. In some examples, other network microphone devices (e.g., NMDs 103b and 103f) in the vicinity of the source of the voice input (e.g., the user) may also detect the voice input. In such cases, the network microphone devices may arbitrate among themselves to determine which device should capture and / or process the detected voice input. Examples of selection and arbitration among network microphone devices are described, for example, in U.S. Patent Application Publication No. 15 / 438,749, filed February 21, 2017, entitled “Voice Control of a Media Playback System,” which is incorporated herein by reference in its entirety.
[0029] In certain embodiments, a network microphone device may be assigned to a playback device that may not include a network microphone device. For example, NMD 103f may be assigned to nearby playback devices 102i and / or 102l. In a related example, a network microphone device may output audio through the playback device to which it is assigned. Further details regarding associating network microphone devices and playback devices as named or default devices are described, for example, in the previously referenced U.S. patent application Ser. No. 15 / 438,749.
[0030] Further aspects of the different components of the exemplary media playback system 100 and how the different components interact to provide a media experience to a user are described in the following sections. Although the description herein may generally refer to the media playback system 100, the techniques described herein are not limited to other applications, such as within a home environment as shown in FIG. 1. For example, the techniques described herein may be useful in other home environment configurations that include more or fewer playback devices, network microphone devices, and / or controller devices 102-104. In addition, the techniques described herein may be useful in environments where multi-zone audio is desired, such as commercial environments such as restaurants, malls, or airports, vehicles such as sports utility vehicles (SUVs), buses or cars, ships or boards, airplanes, etc.
[0031] a. Exemplary Playback Devices and Network Microphone Devices 2A is a functional block diagram illustrating one aspect of a selected one of the playback devices 102 shown in FIG. 1. As shown, such a playback device may include a processor 212, software components 214, memory 216, audio processing components 218, an audio amplifier 220, a speaker 222, and a network interface 230 including a wireless interface 232 and a wired interface 234. In some embodiments, the playback device may not include the speaker 222, and instead may include a speaker interface for connecting the playback device to external speakers. In certain embodiments, the playback device may not include either the speaker 222 or the audio amplifier 222, and instead may include an audio interface for connecting the playback device to an external audio amplifier or audiovisual receiver.
[0032] The playback device may further include a user interface 236. The user interface 236 may facilitate user interaction independent of or in conjunction with one or more of the controller devices 104. In various embodiments, the user interface 236 includes one or more actual buttons and / or a touch-sensitive screen and / or a graphical interface provided on a surface for the user to provide direct input, among other possibilities. The user interface 236 may further include one or more of a light source and a speaker to provide visual and / or audible feedback to the user.
[0033] In some embodiments, the processor 212 may be a clocked computer component configured to process input data based on instructions stored in the memory 216. The memory 216 may be a tangible computer-readable medium configured to store instructions executable by the processor 212. For example, the memory 216 may be a data storage that may include one or more of the software components 214 executable by the processor 212 to perform a function. In one example, the function may involve the playback device reading audio data from an audio source or another playback device. In another example, the function may involve the playback device transmitting audio data to another device on a network. In yet another example, the function may involve pairing the playback device with one or more playback devices to create a multi-channel audio environment.
[0034] One feature may involve a playback device synchronizing the playback of audio content with one or more other playback devices. During synchronized playback, a listener may not notice a delay in the playback of audio content between the synchronized playback devices. U.S. Patent No. 8,234,395, entitled "System and Method for Synchronizing Operation Among Multiple Independently Clocked Digital Data Processing Devices," issued April 4, 2004, is incorporated herein by reference in its entirety, and provides more detailed examples of synchronizing audio playback between playback devices.
[0035] The audio processing component 218 may include one or more digital-to-analog converters (DACs), audio processing components, audio enhancement components, digital signal processors (DSPs), or the like. In some embodiments, one or more of the audio processing components 218 may be subcomponents of the processor 212. In one example, audio content may be processed and / or purposefully altered by the audio processing component 218 to generate an audio signal. The generated audio signal may then be sent for amplification by the audio amplifier 210 and played through the speakers 212. In particular, the audio amplifier 210 may include a device configured to amplify the audio signal to a level capable of driving one or more of the speakers 212. The speakers 212 may comprise a complete speaker system including a separate transducer (e.g., a "driver") or a housing that contains one or more drivers. The particular drivers included in the speakers 212 may include, for example, a subwoofer (e.g., for low frequencies), a mid-range driver (e.g., for mid-frequency frequencies), and / or a tweeter (for high frequencies). In some cases, each transducer of the one or more speakers 212 may be driven by a corresponding individual audio amplifier of the audio amplifiers 210. In addition to generating analog signals for playback, the audio processing component 208 may be configured to process audio content and send the audio content for playback to one or more other playback devices.
[0036] Audio content to be processed and / or played by the playback device may be received from an external source, for example, via an audio line-in input connection (eg, an auto-detecting 3.5 mm audio line-in connection) or via a network interface 230.
[0037] The network interface 230 may be configured to facilitate data flow between the playback device and one or more other devices over a data network. In this manner, the playback device may be configured to receive audio content over a data network from one or more other playback devices in communication with the playback device, from network devices in a local area network, or from an audio content source on a wide area network, such as the Internet. In one example, audio content and other signals sent and received by the playback device may be transmitted in the form of digital packet data that includes an Internet Protocol (IP)-based source address and an IP-based destination address. In such a case, the network interface 230 may be configured to parse the digital packet data so that data destined for the playback device can be appropriately received and processed by the playback device.
[0038] As shown, the network interface 230 may include a wireless interface 232 and a wired interface 234. The wireless interface 232 may provide a network interface function for the playback device to wirelessly communicate with other devices (e.g., other playback devices, speakers, receivers, network devices, control devices in a data network associated with the playback device) based on a communication protocol (e.g., any of the wireless standards (standards) including wireless standards IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G mobile communication standards, etc.). The wired interface 234 may provide a network interface function for the playback device to communicate via a wired connection with other devices based on a communication protocol (e.g., IEEE 802.3). Although the network interface 230 shown in FIG. 2 includes both the wireless interface 232 and the wired interface 234, the network interface 230 may include only a wireless interface or only a wired interface in some embodiments.
[0039] As described above, the playback device may include a network microphone device, such as one of the NMDs 103 shown in FIG. 1. The network microphone device may share some or all of the components of the playback device, such as the processor 212, memory 216, and microphone 224. In another example, the network microphone device includes components that are dedicated only to aspects of the network microphone device's operation. For example, the network microphone device may include a far-field microphone and / or voice processing components that the playback device may not include. In another example, the network microphone device may include a touch-sensitive button to turn the microphone on and off. In yet another example, the network microphone device may be a separate device as described above. FIG. 2B is an isometric view of an example playback device 202 incorporating a network microphone device. The playback device 202 has a control area 237 on the top of the device for turning the microphone on and off. The control area 237 is next to another area 239 on the top of the device for controlling playback.
[0040] By way of example, Sonos Incorporated currently offers (or has offered) playback devices including "PLAY:1", "PLAY:3", "PLAY:5", "PLAYBAR", "CONNECT:AMP", "CONNECT", and "SUB". Any other past, present, and / or future playback devices may additionally or alternatively be implemented and used with the playback device embodiments disclosed herein. Furthermore, it is understood that the playback device is not limited to the specific example shown in FIG. 2A or the Sonos products provided. For example, the playback device may include wired or wireless headphones. In another example, the playback device may include or interact with a docking station for a personal mobile media playback device. In yet another example, the playback device may be integrated with another device or component, such as a television, a lighting fixture, or some other device for use indoors or outdoors.
[0041] b. Exemplary Playback Device Configurations 3A-3E show examples of configurations of playback devices in zones and zone groups. First, referring to FIG. 3E, as one example, a single playback device may belong to one zone. For example, the playback device 102c in "Balcony" may belong to "Zone A". In some embodiments described below, multiple playback devices may be "bonded" to form a "bonded pair" that together form one zone. For example, the playback device 102f named "Nook" in FIG. 1 may be bonded to the playback device 102g named "Wall" to form "Zone B". The bonded playback devices may have different playback roles (e.g., channel roles). In another embodiment described below, multiple playback devices may be merged to form one zone. For example, the playback device 102d named "Office" may be merged with the playback device 102m named "Window" to form "Zone C". The merged playback devices 102d and 102m may not be assigned different playback roles. That is, merged playback devices 102d and 102m may each play the audio content as if they were not merged, except that they play the audio content synchronously.
[0042] Each zone of the media playback system 100 may be provided for control as a single user interface (UI) entity. For example, "Zone A" may be provided as a single entity named "Balcony." "Zone C" may be provided as a single entity named "Office." "Zone B" may be provided as a single entity named "Shelf."
[0043] In various embodiments, a zone may take on the name of one of the playback devices that belong to that zone. For example, "Zone C" may take on the name of the "Office" device 102d (as shown). In another example, "Zone C" may take on the name of the "Window" device 102m. In a further example, "Zone C" may use a name that is some combination of the "Office" device 102d and the "Window" device 102m. The name chosen may be selected by the user. In some embodiments, a zone may be given a name other than the devices that belong to that zone. For example, "Zone B" may be named "Shelf", but none of the devices in "Zone B" have this name.
[0044] The combined playback devices may have different playback responsibilities, such as specific audio channel responsibilities. For example, as shown in FIG. 3A, "nook" and "wall" devices 102f and 102g may be combined to create or enhance a stereo effect for audio content. In this example, the "nook" playback device 102f may be configured to play a left channel audio component, and the "wall" playback device 102g may be configured to play a right channel audio component. In some embodiments, such a stereo combination may be referred to as "pairing."
[0045] In addition, the combined playback devices may each have additional and / or different speaker drivers. As shown in FIG. 3B, a playback device 102b labeled "Front" may be combined with a playback device 102k labeled "Sub". The "Front" device 102b may provide a mid-to-high frequency range, and the "Sub" device 102k may provide low frequencies, for example as a subwoofer. When uncoupled, the "Front" device 102b may provide a full frequency range. As another example, FIG. 3C shows the "Front" and "Sub" devices 102b and 102k further combined as "Right" and "Left" devices 102a and 102j, respectively. In some embodiments, the "Right" and "Left" devices 102a and 102j may form a surround or "satellite" channel of a home theater system. The combined playback devices 102a, 102b, 102j, and 102k may form a single "Zone D" (FIG. 3E).
[0046] The merged playback devices may not have assigned playback responsibilities and each may provide the full range of audio content that each playback device is capable of. Nevertheless, the merged devices may represent a single UI entity (i.e., one zone, as described above). For example, the "office" playback devices 102d and 102m are a single UI entity called "Zone C." In one embodiment, the playback devices 102d and 102m may each synchronously output the full range of audio content that each playback device 102d and 102m is capable of.
[0047] In some embodiments, an independent network microphone device may exist alone in a zone. For example, the NMD 103g of FIG. 1 named "Ceiling" may be "Zone E". A network microphone device may also be combined or merged with another device to form a zone. For example, the NMD device 103f named "Island" may be combined with the playback device 102i "Kitchen" to form together "Zone G", also named "Kitchen". Further details regarding associating network microphone devices and playback devices as named or default devices are described, for example, in the previously referenced U.S. Patent Application Publication No. 15 / 438,749. In some embodiments, an independent network microphone device may not be associated with a zone.
[0048] The zones of individual, combined, and / or merged devices may be grouped to form zone groups. For example, with reference to FIG. 3E, “Zone A” may be grouped with “Zone B” to form a zone group containing the two zones. As another example, “Zone A” may be grouped with one or more other “Zones CI”. “Zones AI” may be grouped and ungrouped in a number of ways. For example, three, four, five, or more (e.g., all) of “Zones AI” may be grouped. Once grouped, individual and / or combined playback devices may play audio in sync with each other, as described in the previously referenced U.S. Pat. No. 8,234,395. Playback devices may be dynamically grouped and ungrouped to form new or different groups to play audio content in sync.
[0049] In various examples, a zone in an environment may be a combination of the default names of the zones in the group or the names of the zones in the zone group, such as "Dining Room + Kitchen" shown in Figure 3E. In some embodiments, a zone group may be given a unique name selected by the user, such as "Nick's Room" also shown in Figure 3E.
[0050] 2A, certain data may be stored as one or more state variables in memory 216 that are periodically updated and used to describe the state of the playback zones, playback devices, and / or their associated zone groups. Memory 216 may also contain data related to the state of other devices in the media system and that is shared between devices from time to time, so that one or more of the devices have the most up-to-date data related to the system.
[0051] In some embodiments, the memory may store instances of various variable types associated with states. The variable instances may be stored with an identifier (e.g., tag) corresponding to the type. For example, a particular identifier may be a first type "a1" that identifies a playback device in a zone, a second type "b1" that identifies a playback device that may be associated with the zone, and a third type "c1" that identifies a zone group to which the zone may belong. As a related example, the identifier associated with "Balcony" in FIG. 1 may indicate that "Balcony" is the only playback device in a particular zone that is not in a zone group. The identifier associated with "Living Room" may indicate that "Living Room" is not grouped with other zones, but includes associated playback devices 102a, 102b, 102j, and 10k. The identifier associated with "Dining Room" may indicate that "Dining Room" is part of a group "Dining Room+Kitchen" and that devices 103f and 102i are associated. The identifier associated with "Kitchen" may indicate the same or similar information by virtue of "Kitchen" being part of the "Dining Room+Kitchen" zone group. Other exemplary zone variables and identifiers are described below.
[0052] In yet another example, the media playback system 100 may have variables or identifiers that represent other associations of zones and zone groups, such as identifiers associated with "areas" shown in FIG. 3. An area may participate in a zone group and / or a cluster of zones that do not belong to a zone group. For example, FIG. 3E shows a first area named "Front Area" and a second area named "Back Area." The "Front Area" includes the following zones and zone groups: "Balcony," "Living Room," "Dining Room," "Kitchen," and "Bathroom." The "Back Area" includes the following zones and zone groups: "Bathroom," "Nick's Room," "Bedroom," and "Office." In one aspect, an "area" may be used to trigger a zone group and / or a cluster of zones that share one or more zones and / or zone groups of another cluster. In another aspect, this is distinct from a zone group that does not share a zone with another zone group. Further examples of techniques for implementing "areas" are described, for example, in U.S. patent application Ser. No. 15 / 682,506, entitled "Name-Based Room Association," filed Aug. 21, 2017, and U.S. Patent No. 8,483,853, entitled "Multi-Zone Media System Control and Group Creation Operation," filed Sep. 11, 2007, which are incorporated herein by reference in their entireties. In some embodiments, media playback system 100 may not implement "areas," in which case the system may not store variables associated with "areas."
[0053] The memory 216 may also be configured to store other data. Such data may relate to audio sources accessible by the playback device or playback queues to which the playback device (or another playback device) may be associated. In an embodiment described below, the memory 216 is configured to store a set of command data for selecting a particular VAS, such as the first VAS 160, when processing a voice input.
[0054] In operation, one or more playback zones in the environment of FIG. 1 may each be playing different audio content. For example, a user may listen to hip hop music played by playback device 102c while grilling in the "balcony" zone, while another user may listen to classical music played by playback device 102i while preparing a meal in the "kitchen" zone. In another example, a playback zone may play the same audio content in synchronization with another playback zone. For example, when a user is in the "office" zone, playback device 102d in the "office" zone may play the same hip hop music as the music being played by playback device 102c in the "balcony" zone. In such a case, playback devices 102c and 102d are playing hip hop music in synchronization, allowing a user to seamlessly (or at least nearly seamlessly) enjoy the audio content played out-loud as they move between different playback zones. Synchronization between playback zones may be performed in a manner similar to synchronization between playback devices as described in the aforementioned U.S. Pat. No. 8,234,395.
[0055] As alluded to above, the zone configuration of the media playback system 100 may change dynamically. Thus, the media playback system 100 may support many configurations. For example, if a user physically moves one or more playback devices into or out of a zone, the media playback system 100 may be reconfigured to accommodate the change. For example, if a user physically moves playback device 102c from the “balcony” zone to the “office” zone, the “office” zone may include both playback device 102c and playback device 102d from there. In some cases, the user may pair or group the moved playback device 102c with the “office” zone and / or rename the playback devices in the “office” zone, for example, using the controller device 104 and / or voice input. As another example, if one or more playback devices are moved to a particular area in the home environment that does not yet have a playback zone, the moved playback devices may be renamed or associated with a playback zone in the particular area.
[0056] Further, different playback zones of the media playback system 100 may be dynamically combined into zone groups or split into separate playback zones. For example, the “dining room” zone and the “kitchen” zone may be combined into a zone group for a dinner party, so that playback devices 102i and 102l play audio content in a synchronized manner. Meanwhile, the combined playback devices 102 in the “living room” zone may be split into (i) a television zone and (ii) a separate listening zone. The television zone may include the “front” playback device 102b. The listening zone may include the “right”, “left”, and “sub” playback devices 102a, 102j, and 102k, which may be grouped, paired, or merged as described above. Such a division of the “living room” zone may allow one user to listen to music in a listening zone in one area of the living room, while another user watches television in another area of the living room space. In a related example, a user may implement either NMD 103a or 103b to control the "living room" zone before it is divided into a television zone and a listening zone. Once divided, the listening zone may be controlled by a user proximate to, for example, NMD 103a, and the television zone may be controlled by a user proximate to, for example, NMD 103b. However, as discussed above, either NMD 103 may be configured to control various playback and other devices of media playback system 100.
[0057] c. Exemplary Controller Device 4 is a functional block diagram illustrating certain aspects of a selected one of the controller devices 104 of the media playback system 100 of FIG. 1. Such a controller device may also be referred to as a controller. The controller device illustrated in FIG. 3 may generally include components similar to certain components of the network devices described above, such as a processor 412, a memory 416, a microphone 424, and a network interface 430. In one example, the controller device may be a dedicated controller device of the media playback system 100. In another example, the controller device may be a network device, such as an iPhone, iPad, or other smartphone, tablet, or network device (e.g., a networked computer such as a PC or Mac), on which controller application software of the media playback system may be installed.
[0058] The controller device's memory 416 may be configured to store controller application software and other data related to the media playback system 100 and users of the system 100. The memory 416 may also be populated with one or more software components 414 executable by the processor 412 to accomplish certain functions, such as facilitating user access, control, and configuration of the media playback system 100. As mentioned above, the controller device communicates with other network devices over a network interface 430, such as a wireless interface.
[0059] In one example, data and information (e.g., state variables, etc.) may be communicated between the controller device and other devices via network interface 430. For example, configurations of playback zones and zone groups for media playback system 100 may be received by the controller device from a playback device, network microphone device, or other network device, or transmitted by the controller device to another playback device or network device, via network interface 406. In some cases, the other network device may be another controller device.
[0060] Playback device control commands, such as volume control and audio playback controls, may also be communicated from the controller device to the playback devices via network interface 430. As described above, configuration changes of media playback system 100 may also be performed by a user using the controller device. Configuration changes may include adding / removing one or more playback devices, adding / removing one or more zones to a zone group, forming combined or merged playback devices, separating one or more playback devices from a combined or merged playback device, etc.
[0061] The controller device user interface 440 may be configured to facilitate user access and control of the media playback system 100 by providing a controller interface, such as controller interfaces 440a and 440b, shown in Figures 4A and 4B, respectively, and collectively referred to as controller interfaces 440. With simultaneous reference to Figures 4A and 4B, the controller interface 440 includes a playback control area 442, a playback zone area 443, a playback status area 444, a playback queue area 446, and a source area 448. The illustrated user interface 400 is only one example of a user interface that may be provided on a network device, such as the controller device shown in Figure 3, accessed by a user to control a media playback system, such as the media playback system 100. Other user interfaces with different formats, styles, and interaction sequences may alternatively be implemented on one or more network devices to provide similar control access to the media playback system.
[0062] Playback control area 442 (FIG. 4A) may include selectable icons (e.g., by touch or with a cursor) that cause playback devices in a selected playback zone or zone group to play or stop, fast forward, rewind, skip next, skip previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit crossfade mode, etc. Playback control area 442 may also include selectable icons to change equalizer settings, playback volume, and various other changes.
[0063] Playback control area 443 (FIG. 4B) may include a representation of the playback zones within media playback system 100. The playback zones area may also include a representation of zone groups, such as the illustrated "dining room + kitchen" zone group. In some embodiments, the graphical representation of the playback zones may be selectable to display additional selectable icons for managing or configuring the playback zones within the media playback system, such as, for example, creating combined zones, creating zone groups, splitting zone groups, renaming zone groups, etc.
[0064] For example, as shown, a "group" icon may be provided in each of the graphical representations of the playback zones. The "group" icon provided in the graphical representation of a particular zone may be selectable to display a selection of one or more other zones of the media playback system to be grouped with the particular zone. Once grouped, the playback devices in the zone grouped with the particular zone are configured to play audio content in synchronization with the playback devices of the particular zone. Similarly, a "group" icon may be provided in the graphical representation of a zone group. In this case, the "group" icon may be selectable to display a selection to deselect one or more zones in the zone group to be removed from the zone group. Other interactions and implementations for grouping and ungrouping zones via a user interface, such as user interface 400, are also possible. The display of the playback zones in playback zone area 443 (FIG. 4B) may be dynamically updated as the configuration of the playback zones or zone groups is changed.
[0065] Playback status area 444 (FIG. 4A) may include a graphical representation of currently played, previously played, or next scheduled to be played in the selected playback zone or zone group. The selectable playback zones or groups may be visually distinguished on the user interface, for example, in playback zone area 443 and / or playback status area 444. The graphical representation may include track title, artist name, album name, album year, track length, and other relevant information that may be useful to a user when controlling the media playback system via user interface 440.
[0066] The playback queue area 446 may include a graphical representation of the audio content in a playback queue associated with a selected playback zone or zone group. In some embodiments, each playback zone or zone group may be associated with a playback queue that includes information corresponding to zero or more audio items to be played by the playback zone or zone group. For example, each audio item in the playback queue may include a Uniform Resource Identifier (URI), Uniform Resource Locator (URL), or other identifier that can be used by a playback device in the playback zone or zone group to locate and / or retrieve the audio item from a local or network audio content source and play it by the playback device.
[0067] In one example, a playlist may be added to the play queue. In this case, information corresponding to each audio item in the playlist may be added to the play queue. In another example, the audio items in the play queue may be saved as a playlist. In yet another example, the play queue may be empty or full but "unused" when the playback device continues to play streaming audio content, e.g., Internet radio, which may be played continuously unless stopped, rather than individual audio items having a play time. In another embodiment, the play queue may contain Internet radio and / or other streaming audio content items and may be "in use" when a playback zone or zone group is playing those items. Other examples are possible.
[0068] When a playback zone or zone group is "grouped" or "ungrouped", the playback queues associated with the affected playback zones or zone groups may be cleared or reassociated. For example, when a first playback zone containing a first playback queue is grouped with a second playback zone containing a second playback queue, the formed zone group may have an associated playback queue that may initially be empty, may contain audio items of the first playback queue (e.g., if the second playback zone is added to the first playback zone), may contain audio items of the second playback queue (e.g., if the first playback zone is added to the second playback zone), or may combine audio items of both the first and second playback queues. If the formed zone group is subsequently ungrouped, the resulting first playback zone may be reassociated with the previous first playback queue, may be associated with a new playback queue that is empty, or may be associated with a new playback queue that contains audio items of the playback queue that was associated with the zone group before the zone group was ungrouped. Similarly, the ungrouped second playback zone may be reassociated with the previous second play queue, may be associated with an empty new play queue, or may be associated with a new play queue containing the audio items in the play queue that were associated with the zone group before the zone group was ungrouped. Other examples are also possible.
[0069] With continued reference to FIGS. 4A and 4B, the graphical representation of the audio content in the play queue area 446 (FIG. 4B) may include track title, artist name, track length, and other relevant information associated with the audio content in the play queue. In some examples, the graphical representation of the audio content may include additional selectable icons that may be selected and moved to allow management and / or editing of the play queue and / or the audio content displayed in the play queue. For example, the displayed audio content may be removed from the play queue, moved to a different position in the play queue, selected to play immediately or after currently playing audio content, or other actions may be performed. A play queue associated with a play zone or zone group may be stored in the memory of one or more playback devices in the play zone or zone group, in the memory of a playback device not in the play zone or zone group, and / or in the memory of other designated devices. Playback of such a play queue may involve one or more playback devices playing the media items in the queue, possibly in sequential or random order.
[0070] The source area 448 may include a graphical representation of selectable audio content and selectable voice assistants associated with a corresponding VAS. The VAS may be selectively assigned. In some examples, multiple VASs, such as AMAZON's ALEXA® and another voice service, may be invoked by the same network microphone device. In some embodiments, a user may assign a VAS to one or more network microphone devices. For example, a user may assign a first VAS 160 to one or both of the "Living Room" NMDs 102a and 102b shown in FIG. 1 and a second VAS to the "Kitchen" NMD 103f. Other examples are possible.
[0071] d. Exemplary Audio Content Sources An audio source in source area 448 may be an audio content source from which audio content is retrieved and played in a selected playback zone or zone group. One or more playback devices in a zone or zone group may be configured to retrieve audio content for playback from multiple available audio content sources (e.g., based on the audio content's corresponding URI or URL). In one example, audio content may be retrieved by a playback device directly from a corresponding audio content source (e.g., a line-in connection). In another example, audio content may be provided to the playback device over a network via one or more other playback devices or network devices.
[0072] Exemplary audio content sources may include memory of one or more playback devices in a media playback system, such as the media playback system 100 of FIG. 1, a local music library on one or more networked devices (e.g., a controller device, a network-enabled personal computer, or a network-attached storage (NAS)), a streaming audio service that provides audio content over the Internet (e.g., the cloud), or an audio source connected to the media playback system via a line-in input connection of a playback device or a network device, or other possible system.
[0073] In some embodiments, audio content sources may be periodically added or removed from a media playback system, such as the media playback system 100 of FIG. 1. In some examples, audio item indexing may occur each time one or more audio content sources are added, removed, or updated. Indexing of audio items may include scanning for identifiable audio items in all folders / directories shared over a network, where the network is accessible by playback devices in the media playback system. Indexing of audio items may also include creating or updating an audio content database that includes metadata (e.g., title, artist, album, track length, etc.) and other related information, which may include, for example, a URI or URL for each identifiable audio item found. Other examples for managing and maintaining audio content sources are possible.
[0074] e. Exemplary Network Microphone Device 5A is a functional block diagram illustrating one or more additional features of NMD 103 according to an embodiment of the disclosure. The network microphone device illustrated in FIG. 5A may include components generally similar to certain components of the network microphone devices described above, such as processor 212 (FIG. 1), network interface 230 (FIG. 2A), microphone 224, and memory 216. Although not shown for clarity, the network microphone device may include other components, such as speakers, amplifiers, signal processors, etc., as described above.
[0075] The microphones 224 may be multiple microphones configured to detect sounds in the network microphone device environment. In one example, the microphones 224 may be configured to detect audio from one or more directions from the network microphone device. The microphones 224 may be sensitive to a portion of a frequency range. In one example, a first subset of the microphones 224 may be sensitive to a first frequency range and a second subset of the microphones 224 may be sensitive to a second frequency range. The microphones 224 may further be configured to obtain location information of an audio source (e.g., a voice, an audible sound) and / or may be configured to help filter background noise. Of note, in some embodiments, the microphones 224 may include a single microphone rather than multiple microphones.
[0076] The network microphone device may further include a beamformer component 551, an acoustic echo cancellation (AEC) component 552, a voice activity detector component 553, a wake word detector component 554, a speech-to-text conversion component 555 (e.g., voice-to-text and text-to-voice), and a VAS selector component 556. In various embodiments, one or more of the components 551-556 may be subcomponents of the processor 512.
[0077] The beamforming and AEC components 551 and 552 are configured to detect audio signals and determine aspects of the voice input in the detected audio, such as direction, amplitude, frequency spectrum, etc. For example, the beamforming and AEC components 551 and 552 may be used in a process to determine an approximate distance between a networked microphone device and a user speaking to the networked microphone device. In another example, a networked microphone device may detect the relative proximity of a user to another networked microphone device of a media playback system.
[0078] The voice activity detector component 553 works in close conjunction with the beamforming and AEC components 551 and 552 and is configured to capture sounds from the direction where voice activity is detected. By monitoring metrics that distinguish speech from other sounds, potential speech directions can be identified. Such metrics may include, for example, the energy in the speech band relative to background noise, and the entropy in the speech band, which is a measure of the spectral structure. Speech typically has a lower entropy than most common background noises.
[0079] The wake word detector component 554 is configured to monitor and analyze the received audio to determine whether any wake words are present in the audio. The wake word detector component 554 may analyze the received audio using a wake word detection algorithm. If the wake word detector 554 detects the wake word, it may process a voice input contained in the audio received by the network microphone device. An exemplary wake word detection algorithm accepts audio as input and provides an indication of whether the wake word is present in the audio. Many first-party and third-party wake word detection algorithms are known and commercially available. For example, a voice service operator may create a proprietary algorithm to use on a third-party device. Alternatively, an algorithm may be trained to detect a particular wake word.
[0080] In some embodiments, wake word detector 554 runs multiple wake word detection algorithms simultaneously (or substantially simultaneously) on the received audio. As discussed above, different voice services (e.g., AMAZON's ALEXA®, APPLE's SIRI®, or MICROSOFT's CORTANA®) use different wake words to invoke their respective voice services. To support multiple services, wake word detector 554 may run the wake word detection algorithms of each supported voice service in parallel on the received audio.
[0081] The VAS selector component 556 is configured to detect commands spoken by the user in the voice input. The speech-to-text conversion component 555 can facilitate the processing step by converting the speech in the voice input to text. In some embodiments, the network microphone device may include voice recognition software trained on a particular user or a particular group of users associated with a household. Such voice recognition software may execute voice processing algorithms adapted to a particular voice profile. Adapting to a particular voice profile may require less computationally intensive algorithms than would be the case for a traditional VAS, which typically samples from a large number of users or a variety of requests that do not target a media playback system.
[0082] The VAS selector component 556 is also configured to determine whether certain command criteria are satisfied for a particular command detected in the voice input. The command criteria for a given command in the voice input may be based, for example, on the inclusion of a particular keyword in the voice input. A keyword may be, for example, a word in the voice input that specifies a particular device or group within the media playback system 100. As used herein, the term "keyword" may refer to a single word (e.g., "bedroom") or a group of words (e.g., "living room").
[0083] Additionally or alternatively, the command criteria for a given command may be associated with detection of one or more control state variables and / or zone state variables along with detection of the given command. The control state variables may include, for example, an indicator of the volume level, the cue(s) associated with the device, and the playback state, such as whether the device is playing a cue or paused. The zone state variables may include, for example, an indicator of which zone the playback devices are in, if grouped together. As described in more detail below, the VAS selector component 556 may store a set of command information in memory 216, such as in a data table 590 that includes a list of commands and associated command criteria.
[0084] In some embodiments, one or more of the above-mentioned components 551-556 are operable in conjunction with microphone 224 to detect and store a user's voice profile associated with a user account on media playback system 100. As described below, in some embodiments, the voice profile may be stored as and / or compared to variables stored in a set of command information 590. The voice profile may include timbre or frequency characteristics of the user's voice and / or other unique characteristics of the user, such as those described in previously referenced U.S. patent application Ser. No. 15 / 438,749.
[0085] In some embodiments, one or more of the above-mentioned components 551-556 are operable in conjunction with the microphone array 524 to determine a user's location within the home environment and / or relative to one or more locations of the NMD 103. As described below, the user's location or proximity may be detected and compared to variables stored in the command information 590. Techniques for determining the user's location or proximity may include one or more of the techniques disclosed in previously referenced U.S. patent application Ser. No. 15 / 438,749, U.S. Patent No. 9,084,058, entitled "Sound Field Calibration Using Listener Localization," filed Dec. 29, 2011, and U.S. Patent No. 8,965,033, entitled "Acoustic Optimization," filed Aug. 31, 2012, each of which is incorporated herein by reference in its entirety.
[0086] 5B is a diagram of an example voice input according to aspects of the disclosure. The voice input may be acquired by a network microphone device, such as one or more of the NMDs 103 shown in FIG. 1. The voice input may include a wake word portion 557a and a voice utterance portion 557b (collectively, "voice input 557"). In some embodiments, wake word 557a may be a well-known wake word, such as "Alexa" associated with AMAZON's ALEXA®. In another embodiment, voice input 557 may not include a wake word.
[0087] In some embodiments, the network microphone device may output an audible and / or visible response upon detecting wake word portion 557a. Additionally or alternatively, the network microphone device may output an audible and / or visible response after processing a voice input or a series of voice inputs (e.g., in the case of a multi-turn request).
[0088] The voice utterances 557b may include, for example, one or more spoken commands 558 (individually identified as a first command 558a and a second command 558b) and one or more spoken keywords 559 (individually identified as a first keyword 559a and a second keyword 559b). In one example, the first command 558a may be a command to play music, a particular song, album, playlist, etc. In this example, the keyword 559 may be one or more words indicating one or more zones in which music should be played, such as "living room" and "dining room" as shown in FIG. 1. In some examples, the voice utterances 557b may include other information, such as detected pauses (e.g., periods of no speech), such as intervals between words spoken by the user, as shown in FIG. 5B. The pauses may demarcate the location of another command, keyword, or other information spoken by the user within the voice utterances 557b.
[0089] In some embodiments, media playback system 100 is configured to temporarily reduce the volume of the audio content it is playing while it detects wake word portion 557a. Media playback system 100 may restore the volume after processing voice input 557, as shown in FIG. 5B. Such a process may be referred to as ducking, examples of which are disclosed in previously referenced U.S. patent application Ser. No. 15 / 438,749.
[0090] f. Exemplary Network and Remote Computer Systems Figure 6 is a functional block diagram showing further details of the remote computer 105 of Figure 1. In various embodiments, the remote computer 105 may receive voice input over the WAN 107 from one or more of the NMDs 103 as shown in Figure 1. For illustrative purposes, a selected communication path for the voice input 557 (Figure 5B) is shown with arrows in Figure 6. In one embodiment, the voice input 557 processed by the remote computer 105 may include a voice utterance 557b (Figure 5B). In another embodiment, the processed voice input 557 may include both the voice utterance 557b and the wake word 557a (Figure 5B).
[0091] The remote computer 105 includes a system controller 612 including one or more processors, an intent engine 602, and a memory 616. The memory 616 may be a tangible computer-readable storage medium configured to store instructions executable by the system controller 612, and / or one or more playback devices, network microphone devices, and / or controller devices 102-104.
[0092] The intent engine 662 is configured to process the voice input and determine the intent of the input. In some embodiments, the intent engine 662 may be a subcomponent of the system controller 612. The intent engine 662 may interact with one or more databases to process the voice input, such as one or more VAS databases 664. The VAS databases 664 may reside in the memory 616 or may reside elsewhere, such as in one or more memories of the playback devices, the network microphone devices, and / or the controller devices 102-104. In some embodiments, the VAS databases 664 may be updated for adaptive learning and feedback based on the voice input processing. The VAS databases 664 may store various user data, analytics, catalogs, and other information for NLU-related and / or other processing.
[0093] The remote computer 105 may exchange various feedback information, instructions, and / or associated data with the various playback devices, network microphone devices, and / or controller devices 102-104 of the media playback system 100. Such exchanges may be related to or independent of transmitted messages, including voice input. In some embodiments, the remote computer 105 and the media playback system 100 may exchange communications via communications paths as described herein and / or using a metadata exchange channel as disclosed in previously referenced U.S. patent application Ser. No. 15 / 438,749.
[0094] Processing of the voice input by devices of the media playback system 100 may be performed, at least in part, in parallel with processing of the voice input by the remote computer 105. Additionally, the speech-to-text component 555 of the networked microphone device may convert responses from the remote computer 105 into speech for audible output using one or more speakers.
[0095] According to various embodiments of the present disclosure, the remote computer 105 performs the functions of the first VAS 160 of the media playback system 100. Figure 7A is a schematic diagram of an example network system 700 including the first VAS 160. As shown, the remote computer 105 is connected to the media playback system 100 via the WAN 107 (Figure 1) and / or a LAN 706 connected to the WAN 107. In this manner, the various playback devices, network microphone devices, and controller devices 102-104 of the media playback system 100 may communicate with the remote computer 105 to invoke the functions of the first VAS 160.
[0096] The network system 700 further includes a first remote computer 705a (e.g., a cloud server) and a second remote computer 705b (e.g., a cloud server). The second remote computer 705b may be associated with a media service provider 767, such as SPOTIFY® or PANDORA®. In some embodiments, the second remote computer 705b may communicate directly with the first VAS 160 computer. Additionally or alternatively, the second remote computer 705b may communicate with the media playback system 100 and / or other intervening remote computers.
[0097] The first remote computer 705a may be associated with a second VAS 760. The second VAS 760 may be, for example, a traditional VAS provider associated with AMAZON's ALEXA®, APPLE's SIRI®, MICROSOFT's CORTANA®, or other VAS provider. Although not shown for clarity, the networked computer system 700 may further include a remote computer associated with one or more additional VASs, such as an additional traditional VAS. In such an embodiment, the media playback system 100 may be configured to select the first VAS 160 over the second VAS or other VASs.
[0098] Figure 7B is a message flow diagram illustrating various data interactions in the networked computer system 700 of Figure 7A. The media playback system 100 obtains voice input via a network microphone device (block 771), such as one or more of the NMDs 103 shown in Figure 1. As described below, the media playback system 100 may select an appropriate VAS based on the commands and associated command criteria in the set of command information 590 (blocks 771-774). If a second VAS 760 is selected, the media playback system 100 may communicate one or more messages 781 (e.g., packets) containing the voice input to the second VAS 760 for processing.
[0099] Conversely, if the first VAS 160 is selected, the media playback system 100 communicates one or more messages 782 (e.g., packets) containing the voice input to the VAS 160. The media playback system 100 may communicate other information to the VAS 160 contemporaneously with the messages 782. For example, the media playback system 100 may communicate data over a metadata channel, as described in the previously referenced U.S. patent application Ser. No. 15 / 131,244.
[0100] The first VAS 160 may process the voice input in the message 782 to determine the intent (block 775). Based on the intent, the VAS 160 may send one or more response messages 783 (e.g., packets) to the media playback system 100. In some examples, the response message 783 may include a payload that commands one or more devices of the media playback system 100 to perform instructions (block 776). For example, the instructions may command the media playback system 100 to play media content, group devices, and / or perform other functions described below. Additionally or alternatively, the response message 783 from the VAS 160 may include a payload that includes a request for more information, such as in the case of a multi-turn command.
[0101] In some embodiments, a response message 783 sent from the VAS 160 may instruct the media playback system 100 to request media content, such as audio content, from the media service 767. In other embodiments, the media playback system 100 may request the content without relying on the VAS 160. In either case, the media playback system 100 may exchange messages to receive the content, such as via a media stream 784 that may include, for example, audio content.
[0102] In some embodiments, the media playback system 100 may receive audio content from a line-in interface at a playback device, a network microphone device, or other device, through a LAN via a network interface. Exemplary audio content includes one or more audio tracks, talk shows, films, television programs, podcasts, Internet streaming videos, etc., among many possible forms of audio content. The audio content may accompany video (e.g., as an audio track of a video), or the audio content may be content that does not accompany video.
[0103] In some embodiments, for training and adaptive training and learning, the media playback system 100 and / or the first VAS 160 may use the voice input that results in a successful (or unsuccessful) response from the VAS (blocks 777 and 778). The training and adaptive learning may improve the accuracy of voice processing by the media playback system 100 and / or the first VAS 160. In one example, the intent engine 662 (FIG. 6) may update and maintain training learning data in a VAS database 664 for one or more user accounts associated with the media playback system 100.
[0104] III. EXAMPLES OF METHODS AND SYSTEMS FOR ACTIVATING A VAS As mentioned above, embodiments described herein may involve invoking the first VAS 160. In one aspect, the first VAS 160 may provide enhanced control capabilities for the media playback system 100. In another aspect, the first VAS may provide an improved VAS experience for controlling the media playback system 100 as compared to other VASs, such as conventional VASs as described above.
[0105] In some embodiments, a conventional VAS, such as the second VAS 760 shown in Figure 7B, may be invoked by the media playback system 100 to perform relatively basic controls, such as relatively simple play / pause / skip functions. In some embodiments, the second VAS 760 may provide other services that may not be easily invoked by the first VAS 160. For example, in certain implementations, the conventional VAS may provide voice-based Internet searches that the first VAS may not provide.
[0106] 8 is an example flow diagram of a method 800 for invoking a VAS. Method 800 illustrates an embodiment of a method that may be performed in an operating environment that includes, for example, media playback system 100, or other media playback system configured according to embodiments of the present disclosure. In the example described below, method 800 involves selecting a first VAS 160 over a second VAS 760.
[0107] The method 800 may include transmitting and receiving information between various devices and systems as described herein and / or as disclosed in the previously referenced U.S. Patent Application Serial No. 15 / 438,749. For example, the method may include transmitting and receiving information between one or more of the playback device, the network microphone device, the controller device, and the remote computers 102-104 of the playback system, the remote computer 705b of the media service 667, and / or the remote computer 705a of the second VAS 670. Although the blocks in FIG. 8 are illustrated in sequence, the blocks may also be performed in parallel and / or in a different order than that described herein. Also, various blocks may be combined to reduce the number of blocks, split to increase the number of blocks, and / or removed based on the desired implementation.
[0108] In addition, for the method 800 and other processes and methods disclosed herein, the flow diagram illustrates one possible function and operation of the implementation of the present embodiment. That is, each block may represent a module, segment, or portion of a program code that includes one or more instructions executable by a processor to perform a particular logical function or step in the process. The program code may be stored in any type of computer-readable storage medium, such as a storage device including a disk or a hard drive. The computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a computer-readable storage medium that stores data for a short time, such as a register memory, a processor cache, and a random access memory (RAM). The computer-readable storage medium may also include a non-transitory computer-readable storage medium, such as a secondary or persistent long-term storage device, such as a read-only memory (ROM), an optical or magnetic disk, or a CD-ROM. The computer-readable storage medium may also be any other volatile or non-volatile storage system. The computer-readable storage medium may be, for example, a computer-readable storage medium or a tangible storage device. The computer-readable storage medium may include one or more of the memories described above with reference to the various playback devices, network microphone devices, controller devices, and remote computers. In addition, for method 800 and other processes and methods disclosed herein, FIG. 8 illustrates circuitry that is wired to perform certain logical functions of the processes.
[0109] In some embodiments, method 800 may further involve receiving user input to launch applications, receiving user and user account information, determining system parameters, interacting with music services, and / or interacting with a controller to display, select, and enter system information, etc. In various embodiments, method 800 may incorporate example methods and systems described in patent application Ser. No. 15 / 223,218, filed Jul. 29, 2016, entitled "Voice Control of a Media Playback System," which is incorporated herein by reference in its entirety.
[0110] a. storing in memory a set of command information including a list of commands and criteria associated with the commands; At block 801, the method 800 involves storing a set of command information, such as the set of command information 590 stored in the memory 216 of the networked microphone device. With reference to FIG. 8B, an exemplary set of command information 890 may include a list of commands 892. The set of command information 890 may be a data table or other data structure. The set of command information 890 may be stored, for example, in memory of one or more of the playback device, the controller device, the networked microphone device, and / or the remote computer 102-105. In some embodiments, the set of command information 890 may be accessible via a metadata exchange channel and / or any other communication path between the media playback system and the remote computer system.
[0111] In the illustrated example, the set of commands 892 includes first through nth "commands." By way of example, the first "command" may be a command to start playback, such as a user saying "play music." The second "command" may be a control command, such as a transport control command to pause, resume, skip playback, etc. For example, the second command may be a command involving a user requesting to skip to the next track of a song. The third "command" may be a zone target command, such as grouping, combining, and merging playback devices. For example, the third command may be a command relating to a user requesting to group the "living room" and the "dining room."
[0112] The commands described herein are examples and other commands are possible. For example, Figures 9A-9C show tables of additional exemplary start playback, control, and zone target commands. As a further example, the commands may include query commands. Query commands may include queries by a user, such as to inquire about currently playing audio. For example, a user may utter the query command "tell me what's playing in the living room."
[0113] As further shown in FIG. 8B, the commands 892 are associated with command criteria that are also stored within the set of command data 890. For example, a first "command" is associated with one or more first command "criteria_1", a second "command" is associated with one or more second command "criteria_2", and a third "command" is associated with one or more third command "criteria_3". The command criteria may involve a determination associated with a particular variable instance. The variable instance may be stored with an identifier (e.g., a tag) that may or may not be associated with a user account. The variable instances may be updated continuously, periodically, or irregularly to include new custom names that are added or deleted by the user or associated with the user account. The custom name may be any name provided by the user and may or may not already exist in the database.
[0114] A variable instance may be present in a keyword in a voice input, may be referenced as a name and / or value stored in a state table, and / or may be dynamically stored and modified in a state table via one or more playback devices, network microphone devices, controller devices, and / or remote computers 102-105. Exemplary variable instances may include zone variable instances, control state variable instances, target variable instances, and other variable instances. Zone variable instances may relate to identifiers representative of zones, zone groups, playback devices, network microphone devices, combined states, areas, etc., including those described above. Control state variables may relate to the current control states of individual playback and network microphone devices, and / or multiple devices, such as information indicating which device is playing music, the volume of the device, cues stored on the device, etc. Target variable instances may relate to advanced state information, such as specific control states and / or groups of devices, combined devices, and merged devices. Target variable variables may also correspond to calibration states, such as equalizer settings, of various devices in the media playback system 100.
[0115] Other variable instances are possible. For example, a media variable instance may identify media content, such as audio content (e.g., a particular track, album, artist, playlist, station, or genre of music). In some embodiments, a media variable may be identified in response to searching a database for audio or content desired by a user. Media variables may be present in a voice input, referenced, maintained, and updated in a state table or referenced in a query, as described above. As another example, a particular variable instance may indicate a user's location or proximity within the Home environment, whether a user's voice profile is detected in a given voice input, whether a particular wake word is detected, etc. Variable instances may include custom variable instances.
[0116] In particular embodiments, at least some of the criteria stored in the set of command information 890 may include scalar vectors of variable instances or other sets of such variable instances. For example, "criteria_1" may include a vector identifying zone variables representative of the zones shown in media playback system 100 of FIG. 1. Such a vector may include "balcony, living room, dining room, kitchen, office, bedroom, Nick's room." In one embodiment, "criteria_1" may be satisfied if two or more zone variables in the vector are detected as keywords in the voice input.
[0117] The set of command information 890 may also include other information, such as user specific information 894 and custom information 896. The user specific information 894 may be associated with a user account and / or a household identifier (HHI). The custom information 896 may include custom variables, such as a custom zone name, a custom playlist, and / or a custom playlist name, for example. For example, "Nick's Favorites" may be a custom playlist with a custom name created by the user.
[0118] b. Acquire voice input 8A , at blocks 802 and 803, method 800 involves monitoring and detecting a wake word in the voice input. For example, media playback system 100 may analyze received audio representative of the voice input to determine whether the wake word is included. Media playback system 100 may analyze the received audio using one or more wake word detection algorithms, such as the wake word detection component as described above.
[0119] At block 804, method 800 involves acquiring voice input following detection of the wake word at blocks 802 and 803. In various embodiments, the voice input may be acquired via one or more of the NMDs 103 of playback system 100. As used herein, the terms "acquiring" or "acquiring" may refer to a process that includes recording at least a portion of the voice input, such as a voice utterance following a wake word. In some embodiments, the acquired voice input may include the wake word. In certain embodiments described below, the terms "acquiring" or "acquiring" may also refer to recording at least a portion of the voice input and converting the voice input into a particular format, such as text, for example, using speech-to-text conversion.
[0120] c. detecting one or more commands in the captured voice input; At blocks 805 and 806, the method 800 involves detecting one or more commands 892 (FIG. 8B) in the voice input captured at block 804. In various embodiments, the method 800 may detect the commands by analyzing the voice input to determine if one of the commands 892 has a grammar that matches the grammar found in the captured voice input. In this manner, the method 800 may use a grammar match to detect the intent of the command in the voice input. The grammar that matches may be a word, a group of words, a phrase, and the like. In one exemplary command, the user may say, "Play The Beatles on the balcony and in the living room." In this example, the method 800 may recognize that the grammar of "play" matches the grammar of a first start play "command" in the set of command information 890. Additionally, the method 800 may recognize "The Beatles" as a media variable and "balcony" and "living room" as zone variables. Thus, command phrasing may be indicated according to variable instances such as: "Play (media variable) with (first zone variable) and (second zone variable)." Similar commands may include "Play (media variable) with (first zone variable) and (second zone variable)." As explained below, "play" may be synonymous with "play."
[0121] In some embodiments, a user may utter a command with or without a zone variable instance. In one example, a user may provide a voice input simply uttering "Play something by the Beatles." In such a case, method 800 may determine the intent to "play something by the Beatles" in a default zone. In another example, method 800 may determine the intent to "play something by the Beatles" in one or more playback devices based on other command criteria that are met for the command, such as if the user's presence is detected in a particular zone when the user requests to play the Beatles. For example, media playback system 100 may play something by the Beatles in the "living room" zone shown in FIG. 1 if the voice input is detected by the "right" playback device 102a located in this zone.
[0122] Another example command may be a "Play Next" command that adds the selected media content to the top of the queue and causes it to be played next in a zone. An example phrasing of this command may be "Play (media variable) next."
[0123] Another example of a command may be a move or transfer command to move or transfer currently playing music and / or a zone's playback queue from one zone to another. For example, a user may utter a voice input of "move music to (zone variable)," and the command words "move" or "transfer" may correspond to an intent to move the playback state to another zone. As a related example, an intent to move music may correspond to two media playback system commands. The two commands may group a first zone with a second zone and then separate the second zone from the group, effectively transferring the state of the second zone to the first zone.
[0124] The intent of commands and variable instances that may be detected in the voice input may be based on a number of predefined vocabulary expressions that may be associated with the user's intent (e.g., play, pause, add to queue, group, and other transfer controls, as well as controls available, for example, via the control device 104). In some embodiments, the processing of commands and associated variable instances may be based on predefined "slots" within the vocabulary in which the commands and variables are expected to be located. In these and other implementations, the set of words and vocabulary used to determine the user's intent may be updated in response to user customizations and preferences, feedback, and adaptive learning, as described above.
[0125] In some embodiments, different words, vocabulary words, and / or phrases used in a command may relate to the same intent. For example, the inclusion of the command words "play," "listen," or "hear" in the voice input may correspond to synonyms that reflect the same intent that the media playback system play the media content.
[0126] 9A-9C show further examples of synonyms. For example, commands on the left side of table 900 may have certain synonyms shown on the right side of the table. With reference to FIG. 9A, for example, the command "play" in the left column has the same intent as synonymous phrases in the right column including "break it down," "let's jam," and "bust it." In various embodiments, commands and synonyms in table 900 may be added, removed, or edited. For example, commands and synonyms may be added, removed, or edited in response to user customization and preferences, feedback, training, and adaptive learning, as described above. FIGS. 9B and 9C show exemplary synonyms associated with control and zone targets, respectively.
[0127] In some embodiments, variable instances may have synonyms defined in a manner similar to the synonyms of commands. For example, a "balcony" zone variable of media playback system 100 may have a synonym "outside" that represents the same zone variable. As another example, a "living room" zone variable may have synonyms "living area," "television room," and "family room."
[0128] d. Determine whether one or more commands meet the corresponding criteria in the command information set. 8A and 8B together, at block 807, method 800 involves determining that one or more commands detected in block 806 satisfy a command criteria in the set of command information 890. With reference to FIG. 8B, for example, if a first command is detected, method 800 determines whether the first command satisfies "criteria_1," if a second "command" is detected, method 800 determines whether the command satisfies "criteria_2," and so on.
[0129] A command may be compared to multiple sets of command criteria. In some embodiments, a particular set of criteria may be associated with a logical operator. For example, a third "command" is compared to commands "criteria_2" and "criteria_3". These commands are linked by a logical AND operator. Thus, the third "command" requires that two sets of criteria be met. In contrast, an nth "command" is associated with criteria (criteria_x, criterion_y, and criterion_z) linked by a logical OR operator. In this case, the nth "command" needs to meet only one of the sets of command criteria for this command. Various combinations of logical operators, including the XOR operator, are possible to determine whether a command meets a particular command criteria.
[0130] In some embodiments, the command criteria may determine whether the voice input includes one or more commands. For example, a voice input with the command "play on (media variable)" may be accompanied by a second command "also play on (zone variable)." In this example, media playback system 100 may recognize "play" as one command and "also play" as a command criterion that is satisfied by the inclusion of the later command. In some embodiments, the commands in the above examples may correspond to a grouping intent when spoken together in the same voice input.
[0131] In a similar embodiment, the voice input may include two commands or phrases spoken in succession. Method 800 may recognize that such consecutive commands or phrases may be related. For example, a user may provide the voice input "play some classical music" followed by "in the living room" and "in the dining room," a command that infers a grouping of playback devices in the "living room" and "dining room."
[0132] In some embodiments, media playback system 100 may detect pauses of limited length (e.g., one or two seconds) as it processes words or phrases in sequence. In some embodiments, the pauses may be intentionally made by the user to separate commands and phrases to facilitate voice processing of a relatively long series of commands and information. The pauses may be of a predetermined duration sufficient to capture the series of commands and information without causing media playback system 100 to resume monitoring for the wake word at block 802. In one aspect, the user may use such pauses to execute multiple commands without repeating the wake word for each command desired to be executed.
[0133] e. In response to the determining step, select the first VAS, discard the selection of the other VAS, and process the one or more commands via the first VAS. Commands that meet certain predefined command criteria may cause the media playback system 100 to activate a first VAS 160, while commands that do not meet the predefined criteria may activate another VAS, or may not activate a VAS at all. The example method 800 involves sending voice input that is determined to meet the command criteria of a given command in the voice input, as shown in blocks 807 and 808, and sending the voice input to another VAS when a given command does not meet the criteria, as shown in block 809.
[0134] At block 810, the method 800 involves receiving and processing a response from the VAS that received the voice input at block 808. In one embodiment, processing the response from the VAS may include processing instructions from the VAS to execute commands in the voice input, such as play, control, zone target, and other commands as described above. In some embodiments, the remote computer may be instructed to initiate or control playback of content associated with media variables that may have been included in the initial voice input or may have been the result of a database search.
[0135] In some embodiments, processing the response at block 810 may cause the media content to be retrieved. In one embodiment, the media variables may be provided to the media playback system 100 as a result of a database search of media content. In some embodiments, the media playback system 100 may retrieve the media content directly from one or more media services. In another embodiment, the VAS may automatically retrieve the media content in conjunction with processing the voice input received at block 800. In various embodiments, the media variables may be communicated through a metadata exchange channel and / or any communication path established with the media playback system 100. Such communication may initiate content streaming, as described above with reference to FIG. 7B.
[0136] In some embodiments, the database search may return results based on the media variables detected in the voice input. For example, the database search may return artists with an album named the same as the media variable, album names that match or are similar to the media variable, tracks named for the media variable, radio stations for the media variable, playlists named for the media variable, streaming service provider identifiers for content related to the media variable, and / or raw speech-to-text results. Using the "American Pie" example, the search results may return the artist "Don McLean," an album named "American Pie," a track named "American Pie," a radio station named "American Pie" (e.g., an identifier for the Pandora radio station for "American Pie"), track identifiers for "American Pie" on music services (e.g., streaming music services such as SPOTIFY® or PANDORA®) (e.g., SPOTIFY® track identifiers, URIs, and / or URLs for "American Pie"), and / or raw speech-to-text results for "American Pie."
[0137] In some embodiments, method 800 may involve updating a play queue stored on the playback device in response to changes to a playlist or play queue stored on the cloud network, such that portions of the play queue match some or all of the playlist or play queue on the cloud network.
[0138] In response to causing an action within media playback system 100, method 800 may involve updating and / or storing information related to the action at block 810. For example, one or more control states, zone states, zone identifiers, or other information may be updated at block 800. Other information that is updated may include, for example, information identifying a particular playback device that is currently playing a particular media item and / or information that a particular media item has been added to a queue stored on a playback device.
[0139] In some embodiments, processing the response at block 810 may lead to a determination that the VAS needs additional information and audibly prompting the user for this information, as shown in blocks 811 and 812. For example, method 800 may prompt the user for additional information when executing a multi-turn command. In such a case, method 800 may return to block 804 to obtain additional voice input.
[0140] Although the methods and systems are described herein with respect to media content (e.g., music content, video content), the methods and systems described herein may be applied to a variety of content having associated audio playable by a media content playback system. For example, previously recorded sounds that may not be part of a music catalog may be played in response to a voice input. One example is a voice input of "What does a nightingale sound like?" The network microphone system's response to this voice input may not be music content with an identifier, but may instead be a short audio clip. The media playback system may receive information (e.g., memory address, link, URL, file) related to playing the short audio clip, and a command of the media playback system to play the short audio clip. Other examples are possible, such as podcasts, news clips, notification sounds, alerts, etc.
[0141] IV. EXEMPLARY IMPLEMENTATIONS FOR VOICE CONTROL OF A MEDIA PLAYBACK SYSTEM 10A-20B are schematic diagrams illustrating various examples of voice inputs processed by media playback system 100 and control interfaces that may indicate the state of media playback system 100 after or before processing the voice input. As described below, command criteria associated with particular voice commands in the voice input may provide enhanced VAS voice control, such as VAS 160 described above. The voice input may be received by one or more of NMDs 103, which may not be incorporated into one of the playback devices 102, as described above.
[0142] Although not shown for clarity, the voice input in the various examples below may be preceded by a wake word, such as AMAZON's ALEXA® or other wake word, as described above. In one aspect, the same wake word may be used to initiate voice capture of the voice input to be sent to a first VAS, such as a conventional VAS, or a second VAS. In such a case, the user making the voice utterance may not be aware that one VAS has been selected in place of another VAS behind the scenes. In certain embodiments, a unique wake word, such as "Hey, SONOS," may be uttered by the user to activate the first VAS without further consideration. In this case, the playback system 100 may avoid the step of determining to select the first VAS in place of another VAS.
[0143] In one aspect, command criteria may be configured to group devices. In some embodiments, such command criteria may initiate playback simultaneously when voice input involves a media variable and / or when the activated device is associated with a playback queue. For example, FIG. 10A shows a user uttering a voice input to NMD 103a of "Play The Beatles in the Living Room and on the Balcony," and the controller interface of FIG. 10B shows the resulting grouping of "Living Room" and "Balcony." In another example, a user may utter a particular track, playlist, mood, or other information to initiate media playback as described herein.
[0144] The voice input in Figure 10A includes the grammatical structure "play (media variable) in (first zone variable) and (second zone variable)." In this example, the command play satisfies the command criteria requiring two or more zone variables as keywords in the voice input. In some embodiments, the "living room" playback devices 102a, 102b, 102j, and 102k may remain in a combined media playback device configuration before and after the utterance of the voice input shown in Figure 10A.
[0145] In some embodiments, the order in which the zone variables are spoken may determine which playback device is designated as the "group head." For example, when a user speaks a voice input that includes the keyword "living room" followed by the keyword "balcony," this order may determine that "living room" is the group head. The group head may be stored as a zone variable in the set of command information 890. The group head may be a handle to refer to a group of playback devices. When a user speaks a voice input that includes the group handle, the media playback system 100 may detect the intent to refer to all devices grouped with "living room." In this manner, when controlling devices collectively, the user does not need to speak the keyword for each zone in the group of devices. In a related embodiment, the user may speak a voice input that changes the group head to another device or zone. For example, the user may change the group head of the "living room" zone to "balcony" (in such a case, the interface may show the group order as "balcony+living room" rather than "living room+balcony").
[0146] In an alternative example, Figure 10C shows a user uttering the voice input "Play The Beatles" but omitting the other keywords in the voice input of Figure 10A. In this example, if the command does not meet any of the criteria in the set of command information 890 as described above, the voice input may be sent to another VAS.
[0147] In another example, the voice input of "play the Beatles" with the above-mentioned keyword omitted may still be sent to the first VAS 160 if other command criteria of the command are satisfied. Such other command criteria may include, for example, criteria involving zone variables, control state variables, target variables, and / or other variables. In one aspect, the variable instance may be the user's proximity (e.g., distance determined by calculation or otherwise) from the network microphone device. For example, the voice input of FIG. 10C may be sent to the first VAS 160 when the user is detected in the vicinity (e.g., based on a predetermined radius r1) of the NMD 103. The determination of the proximity may be based, for example, on the signal strength of the voice input source. In another aspect, the voice input of FIG. 10C may be sent to the first VAS 160 when the user's voice profile is detected, regardless of whether the user's proximity is detected.
[0148] In yet another aspect, proximity and / or other command criteria may facilitate resolution of voice inputs that cannot be easily processed by a conventional VAS. For example, as shown in FIG. 11A, a user uttering a voice input of "Turn up the volume on the balcony" may not be resolved by a conventional VAS because the balcony includes a lighting device 108 with the same name. With reference to FIG. 1, the first VAS 160 may be able to resolve such duplicate device names by determining whether the user is in proximity to the playback device 102c and / or whether "Balcony" is currently playing based on the associated control variables. In a related aspect, the first VAS 160 may determine to turn up the volume of the playback device 102c of the "Balcony" where the user is in proximity, but not the volume of the "Living Room" where the user is not. In such a case, as shown in FIG. 11B, the media playback system 100 may turn up the volume of the "Balcony" but not the volume of the "Living Room".
[0149] Similarly, the first VAS 160 may resolve command overlaps between devices with similar command naming conventions. For example, the “dining room” thermostat 110 shown in FIG. 1 may be programmed by a user to set a particular temperature (e.g., a level between 60 degrees and 85 degrees) when the user speaks a voice input of “set it”; similarly, the “dining room” may be set to a particular volume level (e.g., a volume level between 0 and 100 percent) when the user speaks a voice input of “set it.” In one example, a user uttering a voice input of “set the dining room to 75” can be resolved by the first VAS 160 because the “dining room” zone is currently playing based on the command criteria stored in the set of command information 890. In contrast, a conventional VAS may not be able to determine whether to change the volume level of the “dining room” to 75 or to set the temperature of the “dining room” thermostat to 75.
[0150] In various embodiments, voice input may be processed along with other input from the user via the individual playback devices, network microphone devices, and controller devices 102-104. For example, the user may independently control group volume, individual volumes, playback state, etc., using the soft buttons and controls on the interface shown in FIG. 11B. Additionally, in the example of FIG. 11B, the user may press a soft button labeled "Group" to access another interface for manually grouping and ungrouping devices. In one aspect, providing multiple ways to interact with the media playback system 100 via voice input, controller input, and manual device input may provide a smooth continuum of control and an improved user experience.
[0151] As another grouping / degrouping example, a voice input of "Play Bob Marley on the Balcony" may cause "Balcony" to automatically separate from "Living Room". In such a case, "Balcony" may play Bob Marley and "Living Room" may continue playing The Beatles. Alternatively, "Living Room" may stop playing if the command criteria dictate that "Living Room" is no longer the group head of the group of playback devices. In another embodiment, the command criteria may dictate that the devices do not automatically ungroup in response to a start play command.
[0152] The command criteria may be configured to move or relocate currently playing music and / or a zone's playback queue from one zone to another. For example, as shown in FIG. 12A, a user may utter a voice input of "move music from 'living room' to 'dining room'". As shown in the controller interface of FIG. 12B, the request to move music may move music playing in the "living room" to the "dining room". In a related example, a user may move music to the "dining room" by uttering a voice input of "move music here" directly to the NMD 103f near the "dining room" shown in FIG. 1. In this case, the user does not explicitly refer to the "dining room", but the VAS 160 infers the intent based on the user's proximity to the "dining room". In a related embodiment, if the VAS 160 determines that the NMD 103f is coupled to the playback device 102l in the "dining room", it may determine to move music to the "dining room" rather than to another adjacent room (e.g., the "kitchen"). In another example, the playback system 100 may infer information from metadata of the currently playing content. In one such example, a user may say, "Move Let It Be (or The Beatles) to 'The Dining Room,'" identifying the particular music to be moved to the desired playback zone and / or zone group. In this manner, the media playback system can distinguish between content that may be currently playing and / or in the playback queue in different playback zones and / or zone groups to determine which content to move.
[0153] In yet another example, all devices associated with a group head, such as the "Living Room," may stop playing when music is moved from the group head to the "Dining Room." In a related example, the "Living Room" zone may lose its designation as a group head when the music has been moved.
[0154] The command criteria may be configured to add a device to an existing group using a voice input command. For example, as shown in Figures 13A and 13B, a user may add a "Living Room" zone by uttering the voice input "Add 'Living Room' to 'Dining Room'" to form a group with "Dining Room". In a related embodiment, a user may add "Living Room" by uttering the voice input "Play it here too" directly to the NMD 103a of "Living Room" shown in Figure 1. In this case, the user does not explicitly reference the Living Room in the voice input, but the VAS 160 may infer that the "Living Room" should be added based on the user's proximity. In another example, the command "Add Living Room" may be uttered, assuming that the user is in the "Dining Room" when having this intent. In this case, the Dining Room target may be implied from the room containing the input device.
[0155] In yet another example, the user may indicate in the voice input whether the "living room" or the "dining room" will be the group head, or the VAS 160 may request the user to designate the group head.
[0156] As another example of adding or creating a group, a user may instantiate a group using voice input with keywords associated with a custom zone variable. For example, a user may create the "Front Area" custom zone variable described above. A user may instantiate the "Front Area" group by uttering a voice input such as "Play Van Halen in 'Front Area'," as shown in Figures 14A and 14B. The previous "Dining Room" group of Figure 13B may be replaced in response to the voice input shown in Figure 14A.
[0157] The command criteria may be configured to remove devices from an existing group using a voice input command. For example, a user may utter a voice input of "remove 'Balcony'" to remove "Balcony" from the "Front Area" group as shown in Figures 15A and 15B. As another example, a "Stop / Remove" command for Balcony may do the same thing. As explained above, other exemplary synonyms are possible. In yet another example, assuming the user is on the balcony, the user may directly speak to the NMD 103c for "Balcony" in Figure 1 by uttering "Stop here" or "Stop this room", etc. to achieve the same result.
[0158] The command criteria may be configured to select an audio content source to perform an associated function. For example, FIG. 16A shows a user uttering a voice input "I want to watch TV" into the NMD 103a. In response, the media playback system 100 switches the audio content source from a music source to a television source, as shown in FIG. 16B. In some embodiments, the "living room" may be automatically isolated from other zones by instructing the media playback system 100 to play a television source. For example, in FIG. 16B, the "living room" switches to a television source while Van Halen continues to play in the "dining room" and "kitchen". In some examples, the user may subsequently utter a command to play a television source in other zones of the Home environment by grouping, as described above.
[0159] In a related embodiment, the media playback system 100 may store state information indicating when the "Living Room" is connected to a television source. When the "Living Room" is in this state, command criteria may dictate that voice commands related to the television source may be executed by the VAS, such as the source commands shown in FIG. 9B (e.g., enhance speech, change to quiet mode).
[0160] Command criteria may be configured to couple devices together. For example, FIG. 17A shows a user uttering a voice input "I want to watch TV in the front". In response, the VAS 160 may determine that the "front" playback device 102b of FIG. 1 leaves the "living room" zone to form a TV zone based on the command criteria, as shown in FIG. 16B. In a related example, the user may directly utter a voice input to the NMD 103b of the "front" playback device 102b to separate this device. The remaining coupled devices in the living room, i.e., the "right", "left" and "sub" devices 102a, 102j and 102k, may stop playing music. The control interface may also display these devices as no longer part of the "living room" zone.
[0161] As another example of merging, the user may form a different merging configuration with the remaining devices in the living room area after detaching the "front" playback device 102b. For example, as shown in Figures 18A and 18B, the user may form a listening zone by uttering a voice input of "Play Bob Marley on satellite and sub to form a listening zone." The term "satellite" may be a custom zone variable that refers to the "right" playback device 102a and the "left" playback device 102j. The voice input in Figure 18A also starts the playback of Bob Marley in the newly formed listening zone. As further shown in the controller interface in Figure 18B, in the illustrated example, the merging operation of Figures 17A-18B did not interrupt the playback of Van Halen in the "dining room" and "kitchen."
[0162] Command criteria may be configured to pair / bind devices. For example, FIG. 19A shows a multi-turn command where a user utters the voice input of "Pair 'Dining Room' and 'Kitchen' in stereo." In this example, the VAS instructs one or more of the NMDs 103 to prompt the user, asking whether the "Dining Room" zone should be the right channel. If the user confirms that "Dining Room" is the right channel, the "Kitchen" zone becomes the left channel. If the user indicates that "Dining Room" is not the right channel, then "Dining Room" defaults to the left channel and the "Kitchen" zone becomes the right channel. When combined, one of the "Dining Room" and "Kitchen" may be selected as the group head. The VAS may prompt the user to name the combined devices with a name that includes a proper name, such as "Cocina," as shown in FIG. 19B. The "Cocina" zone may resume playing Van Halen, which may be a transfer from the play queue of either the original "Dining Room" or "Kitchen" zone.
[0163] In a related embodiment, device docking and merging can cause the VAS to initiate multi-turn or other commands to calibrate the playback device, as shown in Figures 20A and 20B. In one example, the VAS 160 may continue the multi-turn command sequence of Figure 19A after pairing the "dining room" and "kitchen" zones. In some embodiments, the command criteria may require detection of a user operating one of the controller devices 103 before initiating calibration. In this manner, the VAS 160 may prepare calibration software, such as SONOS' TRUEPLAY®, for calibration, as shown in Figure 20B.
[0164] VII. Conclusion The above description discloses various exemplary systems, methods, apparatus, and articles of manufacture that include, among other things, components, firmware, and / or software executing on hardware. It is understood that such examples are merely exemplary and should not be considered limiting. For example, it is contemplated that any or all of the firmware, hardware, and / or software aspects or components may be implemented exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Thus, the examples provided are not the only ways to implement such systems, methods, apparatus, and / or articles of manufacture.
[0165] (Feature 1) A method of invoking a first voice assistant service (VAS) in a media playback system, the method comprising: storing in a memory a set of command information including a list of commands and associated command criteria; obtaining a voice input via at least one microphone of a network microphone device; Detecting one or more commands included in the voice input; determining that one or more commands satisfy corresponding command criteria in the set of command information; In response to the determination, select the first VAS and discard the selection of the second VAS; (ii) transmit the voice input to the first VAS; and (iii) after transmitting the voice input, receive a response to the voice input from the first VAS. method.
[0166] (Feature 2) A media playback system includes multiple playback devices; the one or more commands group two or more of the playback devices of the plurality of playback devices and initiate playback of the audio content on the group including the two or more playback devices; The method according to feature 1.
[0167] (Feature 3) The determining step includes detecting that one or more keywords are included in the voice input; the one or more keywords include at least one of: (i) a first keyword associated with one of the two or more playback devices and a second keyword associated with another of the two or more playback devices; and (ii) a group including the two or more playback devices; The method according to feature 2.
[0168] (Feature 4) One of the two or more playback devices includes a network microphone device. The method according to feature 2.
[0169] (Feature 5) One or more commands are directed to a media playback system; The method further includes processing one or more commands via the media playback system based on a response from the first VAS. The method according to feature 1.
[0170] (Feature 6) The one or more commands include at least one of a playback command and a transport control command. The method according to feature 5.
[0171] (Feature 7) The voice input is a first voice input, The method further includes outputting an audible request based on a response from the first VAS. The method according to feature 1.
[0172] (Feature 8) The voice input is a first voice input, The method further includes outputting an audible request to a second voice input based on a response from the first VAS. The method according to feature 1.
[0173] (Feature 9) A media playback system includes a plurality of playback devices, the one or more commands include a command to pair two or more playback devices; the audible request includes a request to assign at least one of the two or more playback devices to the audio channel; the second voice input includes a selection of at least one of two or more playback devices; The method according to feature 8.
[0174] (Feature 10) A media playback system includes one or more playback devices, the audible request includes a request to calibrate an equalizer setting of one or more playback devices; The method according to feature 8.
[0175] The determining step includes detecting a presence of a voice input source. The method according to feature 1.
[0176] Feature 12: Detecting the presence includes detecting a direction from which the voice input is received by the network microphone device from the voice input source. 12. The method according to claim 11.
[0177] 13. The method of claim 1, wherein the detecting the presence includes detecting a distance between the network microphone device and the voice input source. 12. The method according to claim 11.
[0178] The determining step includes detecting use of the controller device. The method according to feature 1.
[0179] Feature 15: The determining step includes detecting a voice profile of the voice input source. The method according to feature 1.
[0180] (Feature 16) The one or more commands are one or more first commands; the determining step includes detecting one or more second commands in the voice input; The method according to feature 1.
[0181] Feature 17: The determining step further includes detecting at least one pause in the voice input between the one or more first commands and the one or more second commands. 17. The method according to claim 16.
[0182] (Feature 18) A network microphone device for a media playback system, comprising: (i) a processor; (ii) at least one microphone; and (iii) a tangible computer-readable storage medium having instructions stored thereon that, when executed by the processor, cause the networked microphone device to perform functions of a media playback system; The features are: (a) storing in a memory a set of command information including a list of commands and associated command criteria; (b) obtaining a voice input via at least one microphone; (c) detecting one or more commands included in the voice input; (d) determining that the one or more commands satisfy corresponding command criteria associated with one or more commands in the set of command information; (e) In response to the decision, (i) select a first voice assistant service (VAS) and abandon the selection of a second VAS; (ii) sending the voice input to a first VAS; (iii) receiving a voice input from the first VAS after transmitting the voice input.
[0183] (Feature 19) A media playback system includes a plurality of playback devices; the one or more commands include a command to group two or more of the playback devices and initiate playback of the audio content on the group that includes the two or more playback devices; 20. The network microphone device according to feature 18.
[0184] The determining step includes detecting inclusion of one or more keywords in the voice input; the one or more keywords include at least one of: (i) a first keyword associated with one of the two or more playback devices and a second keyword associated with another of the two or more playback devices; and (ii) a group including the two or more playback devices; 20. A network microphone device according to feature 19.
[0185] (Feature 21) One of the two or more playback devices includes a network microphone device. 20. A network microphone device according to feature 19.
[0186] The one or more commands are directed to a media playback system; The function further includes processing one or more commands via the media playback system based on a response from the first VAS. 20. The network microphone device according to feature 18.
[0187] (Feature 23) The one or more commands include at least one of a playback command and a transport control command. 23. The network microphone device according to feature 22.
[0188] (Feature 24) The voice input is a first voice input, The function further includes outputting an audible request based on a response from the first VAS. 20. The network microphone device according to feature 18.
[0189] (Feature 25) The voice input is a first voice input, The function further includes outputting an audible request to a second voice input based on a response from the first VAS. 20. The network microphone device according to feature 18.
[0190] (Feature 26) A media playback system includes a plurality of playback devices, the one or more commands include a command to pair two or more playback devices; the audible request includes a request to assign at least one of the two or more playback devices to the audio channel; the second voice input includes a selection of at least one of two or more playback devices; 26. A network microphone device as described in feature 25.
[0191] (Feature 27) A media playback system includes one or more playback devices, the audible request includes a request to calibrate an equalizer setting of one or more playback devices; 26. A network microphone device as described in feature 25.
[0192] 28. The step of determining includes detecting a presence of a voice input source. 20. The network microphone device according to feature 18.
[0193] Feature 29: Detecting the presence includes detecting a direction from which the voice input is received by the network microphone device from the voice input source. 30. The network microphone device according to feature 28.
[0194] 30. The method of claim 30, wherein the detecting the presence includes detecting a distance between the network microphone device and the voice input source. 30. The network microphone device according to feature 28.
[0195] 31. The step of determining includes detecting use of the controller device. 20. The network microphone device according to feature 18.
[0196] 32. The step of determining includes detecting a voice profile of the voice input source. 20. The network microphone device according to feature 18.
[0197] (Feature 33) The one or more commands are one or more first commands; the determining step includes detecting one or more second commands in the voice input. 20. The network microphone device according to feature 18.
[0198] Feature 34: The determining step further includes detecting at least one pause in the voice input between the one or more first commands and the one or more second commands. 34. The network microphone device according to feature 33.
[0199] A method for invoking a first voice assistant service (VAS) in a media playback system, comprising: The method is: (i) storing in a memory a set of command information including a list of commands and associated command criteria; (ii) obtaining a voice input via at least one microphone; (iii) detecting one or more commands in the voice input; (iv) determining that the one or more commands satisfy corresponding command criteria associated with one or more commands in the set of command information; (v) In response to the judgment, (a) select a first voice assistant service (VAS) and abandon the selection of a second VAS; (b) sending the voice input to a first VAS; (c) after transmitting the voice input, receiving a voice input from the first VAS.
[0200] (Feature 36) A media playback system includes a plurality of playback devices, the one or more commands include a command to group two or more playback devices of the plurality of playback devices and initiate playback of the audio content on the group including the two or more playback devices; the determining step includes detecting inclusion of one or more keywords within the voice input; the one or more keywords include at least one of: (i) a first keyword associated with one of the two or more playback devices and a second keyword associated with another of the two or more playback devices; and (ii) a group including the two or more playback devices; 36. The method according to claim 35.
[0201] A tangible, non-transitory computer-readable storage medium having stored therein instructions that, when executed by one or more processors, cause a network microphone device to operate within a media playback system, the instructions comprising: The operation is (i) storing in a memory a set of command information including a list of commands and associated command criteria; (ii) obtaining a voice input via at least one microphone; (iii) detecting one or more commands in the voice input; (iv) determining that the one or more commands satisfy corresponding command criteria associated with one or more commands in the set of command information; (v) In response to the judgment, (a) select a first voice assistant service (VAS) and abandon the selection of a second VAS; (b) sending the voice input to a first VAS; (c) after transmitting the voice input, receiving the voice input from the first VAS.
[0202] The present specification has been broadly described in terms of example environments, systems, procedures, steps, logic blocks, processes, and other symbolic representations, which are analogous to the operation of data processing devices directly or indirectly connected to a network. These process descriptions and representations are commonly used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art. Many specific details have been provided to aid in understanding the present disclosure. However, those skilled in the art will understand that certain embodiments of the present disclosure may be practiced without certain, specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail to avoid unnecessarily obscuring the embodiments. Thus, the scope of the present disclosure is defined by the appended claims rather than by the embodiments described above.
[0203] If any of the appended claims are read to cover implementations solely in software and / or firmware, then one or more of the elements in at least one example are expressly defined herein to include a tangible, non-transitory storage medium, e.g., memory, DVD, CD, Blu-ray®, etc., that stores the software and / or firmware.
Claims
1. storing a set of command information in a memory of a network microphone device of a media playback system including a plurality of playback devices, the set of command information including a list of commands for controlling one or more of the playback devices and command criteria associated with the commands, the command criteria including one or more keywords associated with particular commands; obtaining a first voice input via at least one microphone of the networked microphone device; detecting one or more commands included within the first voice input; in response to determining that one or more keywords in the voice input are not in a set of command information stored in a memory of the network microphone device, determining that the one or more commands do not satisfy a corresponding command criteria associated with the one or more commands in the set of command information; and in response to determining that the one or more commands do not satisfy a corresponding command criteria associated with the one or more commands in a set of command information, causing one or more remote servers of a first voice assistant service (VAS) to process at least one keyword of the voice input.
2. moreover, obtaining a second voice input via the at least one microphone; detecting one or more second commands within the second voice input; in response to detecting inclusion of one or more second commands within the second voice input, processing the second voice input by the first VAS by matching keywords of the second voice input to slots associated with the one or more second commands and identifying the one or more second commands and / or command criteria associated with the one or more second commands; Including, The method of claim 1.
3. the one or more second commands include a command to group two or more of the playback devices; The matching includes determining that the second voice input includes an indication of a zone. The method of claim 2.
4. the keywords relate to metadata of the currently playing media content; The method according to any one of claims 1 to 3.
5. moreover selecting a second VAS and abandoning the selection of the first VAS when the third voice input includes one of a wake word associated with the second VAS or a command not included in the second command information; The method according to any one of claims 1 to 4.
6. one of the plurality of playback devices includes the network microphone device; The method according to any one of claims 1 to 5.
7. the one or more second commands are directed to the media playback system; and processing the one or more commands via the media playback system based on a response from the first VAS. The method according to any one of claims 1 to 6.
8. outputting an audible request based on a response from the first VAS. The method according to any one of claims 1 to 7.
9. and outputting an audible request for a third voice input based on a response from the second VAS. The method according to claim 5.
10. the one or more commands include a command to pair two or more playback devices of the plurality of playback devices; the audible request includes a request to assign at least one of the two or more playback devices to an audio channel; the second voice input includes a selection of at least one of the two or more playback devices; 10. The method according to claim 8 or 9.
11. the audible request includes a request to calibrate an equalizer setting of the one or more playback devices.
10. The method according to claim 8 or 9.
12. detecting at least one pause in the voice input between the one or more first commands and the one or more second commands. The method according to any one of claims 1 to 11.
13. The second VAS is a default VAS, and the first VAS provides improved voice control of the media playback system. The method according to any one of claims 1 to 12.
14. A non-transitory computer readable storage medium comprising instructions which, when executed by one or more processors, cause a network microphone device to perform the method of any one of claims 1-13.
15. One or more microphones; one or more processors; The non-transitory computer readable storage medium according to claim 14, Network microphone device.
Citation Information
Patent Citations
Voice control of a media playback system
WO2017147081A1