Network microphone device with command keyword adjustment

By integrating a command keyword engine in NMDs that necessitates specific playback conditions for command execution, the issue of false detections in conventional wake word engines is mitigated, resulting in improved processing efficiency and user privacy through localized natural language processing.

JP7692965B2Active Publication Date: 2025-06-16SONOS INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023149069
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-06-12
Filing Date
2023-09-14
Publication Date
2025-06-16
Estimated Expiration
2040-06-11

AI Technical Summary

Technical Problem

Conventional wake word engines in network microphone devices (NMDs) are prone to false detections due to acoustically similar words or background audio, leading to unnecessary resource consumption and interruptions in audio playback.

Method used

Implementing a command keyword engine that requires both detection of a command keyword and satisfaction of specific playback conditions before generating a command keyword event, thereby reducing false detections and enhancing user privacy by minimizing cloud-based processing.

Benefits of technology

The solution effectively reduces the occurrence of false command keyword events, leading to faster and more accurate voice input processing, while also enhancing user privacy by localizing natural language processing within the NMD.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692965000001
    Figure 0007692965000001
  • Figure 0007692965000002
    Figure 0007692965000002
  • Figure 0007692965000003
    Figure 0007692965000003
Patent Text Reader

Abstract

To provide a method for avoiding false detection.SOLUTION: A playback device includes a voice assistant service (VAS) wake word engine and a command keyword engine. The playback device detects, with the command keyword engine, a first command keyword and determines whether one or more playback conditions corresponding to the first command keyword are satisfied. Based on (a) detecting the first command keyword and (b) determining that one or more playback conditions corresponding to the first command keyword are satisfied, the playback device executes a first playback command corresponding to the first command keyword. When the playback device detects a wake word in an audio input via the wake word engine, the playback device streams sound data corresponding to at least a portion of the audio input to one or more remote servers associated with the VAS.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Cross - reference to related applications

[0001] This application claims the priority of U.S. Patent Application Publication No. 16 / 439,009, entitled "NETWORK MICROPHONE WITH COMMAND KEYWORD CONDITIONING", filed on June 12, 2019, the entire content of which is incorporated herein by reference.

Technical Field

[0002] This technology relates to consumer goods, and more particularly, to methods, systems, products, functions, services, and other elements related to voice - assisted control of a media playback system or some aspect thereof.

Background Art

[0003] Until 2002 when SONOS, Inc. began developing a new type of playback system, the options for accessing and listening to digital audio with the sound output setting were limited. Subsequently, Sonos filed one of its first patent applications, titled "Method for Synchronizing Audio Playback between Multiple Networked Devices," in 2003 and began offering its first media playback system for sale in 2005. The Sonos wireless home sound system enables people to experience music from multiple sources through one or more networked playback devices. Through a software control application installed on a controller (e.g., smartphone, tablet, computer, voice input device), one can play what one desires in any room with a networked playback device. Media content (e.g., songs, podcasts, video sounds) can be streamed to the playback devices so that each room with a playback device can play different corresponding media content. Additionally, rooms can be grouped together for synchronized playback of the same media content, and / or the same media content can be listened to synchronously in all rooms.

Brief Description of the Drawings

[0004] The features, aspects, and advantages of the technology of this disclosure can be better understood in relation to the following description, the appended claims, and the accompanying drawings.

[0005] The features, aspects, and advantages of the technology of this disclosure can be better understood in relation to the following listed subsequent description, the appended claims, and the accompanying drawings. Those skilled in the art will understand that the features shown in the drawings are for illustrative purposes and that variations including different and / or additional features and arrangements are possible.

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 3E

Figure 4

Figure 5A

Figure 5B

Figure 6

Figure 7A

Figure 7B

Figure 7C

Figure 8

Figure 9A

Figure 9B

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15A

Figure 15B

Figure 15C

Figure 15D

[0006] The drawings are for illustrative purposes of exemplary embodiments, but it should be understood that the present invention is not limited to the arrangements and means shown in the drawings. In the drawings, the same reference numbers identify at least generally similar elements. For ease of explanation of any particular element, the most significant digit of any reference number refers to the figure in which that element is first introduced. For example, element 103a is first introduced in Figure 1A and is described with reference to Figure 1A.

Best Mode for Carrying Out the Invention

[0007] I. Overview The exemplary techniques described herein include a wake word engine configured to detect commands. An exemplary network microphone device (“NMD”) may implement such a wake word engine in parallel with a wake word engine that invokes a voice assistant service (“VAS”). While the VAS wake word engine may be involved in non-command wake words, the command keyword engine is invoked with commands such as “play” or “skip”.

[0008] The network microphone device may be used to facilitate voice control of smart home devices such as wireless audio playback devices, illumination devices, household appliances, and home automation equipment (e.g., thermostats, door locks, etc.). The NMD is typically a network computing device that includes an array of microphones, such as a microphone array, configured to detect sounds present in the environment of the NMD. In some examples, the NMD may be implemented within another device, such as an audio playback device.

[0009] Such voice input to an NMD typically includes a wake word followed by speech containing the user's request. In practice, the wake word is typically a predetermined non-word or phrase used to "activate" the NMD and invoke a particular voice assistant service ("VAS") to interpret the intent of the voice input in the detected sound. For example, among other examples, the user can say the wake word "Alexa" to invoke the AMAZON (registered trademark) VAS, "Okay, Google" to invoke the GOOGLE (registered trademark) VAS, "Hey, Siri" to invoke the APPLE (registered trademark) VAS, and "Hey, Sonos" to invoke the VAS provided by SONOS (registered trademark). In practice, the wake word may also be referred to, for example, as a wake word, trigger word, wake-up word or phrase, and can take the form of any suitable word, combination of words (e.g., a particular phrase), and / or any other kind of audio cue.

[0010] To identify whether the sound detected by the NMD includes voice input containing a particular wake word, the NMD often utilizes a wake word engine that is typically installed in the NMD. The wake word engine may be configured to identify (i.e., "spot" or "detect") a particular wake word in the recorded audio using one or more identification algorithms. Such identification algorithms can include pattern recognition trained to detect frequency and / or time domain patterns generated by speaking the wake word. This wake word identification process is generally referred to as "keyword spotting". In practice, to facilitate keyword spotting, the NMD can buffer the sound detected by the NMD's microphone and then use the wake word engine to process the buffered sound to determine whether the wake word is present in the recorded audio.

[0011] When the wake word engine detects a wake word in the recorded audio, the NMD can determine that a wake word event (i.e., "wake word trigger") indicating that the NMD has detected a sound that the voice input will likely contain has occurred. The occurrence of a wake word event typically causes the NMD to perform additional processing related to the detected sound. In the VAS wake word engine, these additional processes can include, among other possible additional processes such as outputting an alert (e.g., an audible chime and / or a light indicator) indicating that the wake word has been identified, extracting the detected sound data from the buffer. Extracting the detected sound can include reading and packaging the stream of the detected sound according to a specific format and sending the packaged sound data to the appropriate VAS for interpretation.

[0012] Next, the VAS corresponding to the wake word identified by the wake word engine receives the sound data transmitted from the NMD via the communication network. The VAS conventionally takes the form of a remote service implemented using one or more cloud servers configured to process voice inputs (e.g., AMAZON's ALEXA, APPLE's SIRI, MICROSOFT's CORTANA, Google's ASSISTANT, etc.). In some cases, specific components and functions of the VAS may be distributed across local and remote devices.

[0013] When the VAS receives the detected sound data, it processes the data, which includes identifying the voice input and determining the intent of the words captured in the voice input. The VAS can then return a response to the NMD with some instruction according to the determined intent. Based on that instruction, the NMD can cause one or more smart devices to perform an action. For example, among other examples, the NMD may cause a playback device to play a specific song or turn an illumination device on / off according to an instruction from the VAS. In some cases, the NMD, or a media system equipped with the NMD (e.g., a media playback system having a playback device equipped with the NMD), may be configured to interact with multiple VASs. In practice, the NMD can select one VAS over another based on a specific wake word identified in the sound detected by the NMD.

[0014] One problem with conventional wake word engines is that they are prone to false detections caused by the triggering of "false wake words". False detections in the context of the NMD generally refer to detected sound inputs that incorrectly call the VAS. In the VAS wake word engine, false detections can call the VAS even when there is no user actually trying to say the wake word to the NMD.

[0015] For example, when the wake word engine identifies a detected sound wake word from the audio being played in the environment of the NMD (e.g., music, podcasts, etc.), false detections may occur. This output audio may be played from a playback device near the NMD or by the NMD itself. For example, when the voice of the AMAZON ALEXA service for commercial advertisements is output near the NMD, the word "Alexa" during the commercial may trigger a false detection. The words or phrases in the output audio that cause false detections may be referred to herein as "false wake words".

[0016] In other examples, words that are acoustically similar to the actual wake word can cause false detections. For example, when the voice of a LEXUS (registered trademark) automobile is emitted as a commercial in the vicinity of an NMD, the word "レクサス" (Lexus) can be a false wake word that causes a false detection because it is acoustically similar to "アレクサ" (Alexa). As another example, false detections can occur when a person speaks a VAS wake word or an acoustically similar word during a conversation.

[0017] The occurrence of false detections is not desirable. This is because it can lead to negative results, among other things, such as consuming additional resources in the NMD or interrupting the audio playback. In some cases of NMDs, such as the AMAZON FIRETV remote or the APPLE TV remote, false detections can be avoided by requiring a button press to invoke the VAS. In practice, the impact of false detections generated by the VAS wake word engine is often partially mitigated by the VAS that processes the detected sound data and determines that the detected sound data does not contain recognizable voice input.

[0018] In contrast to a given one-time-use wake word that invokes VAS, a keyword that invokes a command (referred to herein as a "command keyword") can be a word or combination of words (e.g., phrase) that functions as the command itself, such as a play command. In some implementations, a command keyword can function as both a wake word and the command itself. That is, when the command keyword engine detects a command keyword in the recorded audio, the NMD can determine that a command keyword event has occurred and execute a command corresponding to the detected keyword in response. For example, based on the detection of the command keyword "pause", the NMD pauses playback. One advantage of the command keyword engine is that the recorded audio does not necessarily have to be sent to VAS for processing, which can result in, among other possible advantages, a faster response to voice input and enhanced user privacy. In some implementations described below, a detected command keyword event can trigger one or more subsequent actions, such as local natural language processing of the voice input. In some implementations, a command keyword event can be one of one or more other conditions that must be detected before such an action is triggered.

[0019] According to the exemplary techniques described herein, after detecting a command keyword, the exemplary NMD can generate a command keyword event (and execute a command corresponding to the detected command keyword) only if certain conditions corresponding to the detected command keyword are met. For example, after detecting the command keyword "skip", the exemplary NMD generates a command keyword event (and skips to the next track) only if certain playback conditions indicating that the skip should be executed are met. These playback conditions can include, for example, (i) a first state where a media item is being played, (ii) a second state where a queue is active, and (iii) a third state where the queue contains a media item that follows the media item being played. If any of these conditions are not met, no command keyword event is generated (and no skip is executed).

[0020] Before generating a command keyword event, the occurrence rate of false detection can be reduced by requiring both (a) detection of the command keyword and (b) certain conditions corresponding to the detected command keyword. For example, when playing TV audio, a dialog or other TV audio cannot generate a false detection for the command keyword "skip" because the TV audio input is active (and not a queue). Further, since the NMD generates a wake word event if conditions regarding the state of the control device gate are met, the command keyword is kept in a constantly receivable state (instead of requiring a button press to put the NMD in a voice input receivable state).

[0021] Aspects for adjusting keyword events may be applicable to VAS wake word engines and other conventional non-VAS wake word engines. For example, such adjustments can enable other practical wake word engines in addition to command keyword engines that might otherwise be prone to false detections. For example, NMD can include a streaming audio service wake word engine that supports specific wake words unique to a streaming audio service. For example, after detecting a wake word for a streaming audio service, an exemplary NMD generates a wake word event for the streaming audio service only if certain streaming audio service conditions are met. These playback conditions can include, among other examples, (i) an active subscription to the streaming audio service, and (ii) an audio track from the streaming audio service in the queue.

[0022] Furthermore, a command keyword may be a single word or a phrase. A phrase generally includes more syllables, which generally makes the command keyword more specific and easier to identify by the command keyword engine. Thus, in some cases, a command keyword that is a phrase may have a lower tendency for false detection. Additionally, by using a phrase, more intent can be incorporated into the command keyword. For example, a command keyword of "skip forward" signals that the skip should advance to the next track in the queue rather than return to the previous track.

[0023] Furthermore, the NMD can include a local natural language unit (NLU). In contrast to an NLU implemented on one or more cloud servers that can recognize a wide variety of voice inputs, an exemplary local NLU can recognize a relatively small library of keywords (e.g., 10,000 words and phrases), thereby facilitating a practical implementation for the NMD. When a command keyword engine generates a command keyword event after detecting the command keyword of a voice input, the local NLU can process the speech utterance part of the voice input, search for keywords from the library, and determine the intent from the found keywords.

[0024] If the speech utterance part of the voice input includes at least one keyword from the library, the NMD can execute the command corresponding to the command keyword according to one or more parameters corresponding to the at least one keyword. In other words, the keyword can modify or customize the command corresponding to the command keyword. For example, the command keyword engine may be configured to detect "play" as the command keyword, and the local NLU library can include the phrase "low volume". Then, when the user says "play music at low volume" as a voice input, the command keyword engine generates a command keyword event for "play" and uses the keyword "low volume" as a parameter for the "play" command. Thus, the NMD not only plays based on this voice input but also lowers the volume.

[0025] Exemplary techniques include customizing keywords in a library for users of a media playback system. For example, NMD can populate (fill) a library using names (such as zone names, smart device names, user names) set in the media playback system. Further, NMD can populate (gather) names such as favorite playlists, Internet radio stations, etc. into a local NLU library. Such customization enables the local NLU to more efficiently assist the user with voice commands. Such customization can also be advantageous as it can limit the size of the local NLU library.

[0026] One possible advantage of local NLU is improved privacy. By processing voice utterances locally, the user can avoid sending voice recordings to the cloud (e.g., to the servers of a voice assistant service). Further, in some implementations, NMD can use a local area network to discover playback devices and / or smart devices connected to the network, thereby avoiding providing this data to the cloud. Also, the user's preferences and customizations can remain local to the NMD within the home and perhaps use the cloud only as an optional backup. Other advantages are similarly possible.

[0027] As described above, the exemplary techniques were related to command keywords. The first exemplary implementation includes a network interface, one or more processors, at least one microphone configured to detect sound, at least one speaker, and receives input sound data representing the sound detected by the at least one microphone, and a wake word engine configured to generate a wake word event for a voice assistant service (VAS) when the wake word engine detects a VAS wake word in the input sound data. When a VAS wake word event is generated, the wake word engine streams sound data representing the sound detected by the at least one microphone to one or more servers of the voice assistant service. A command keyword engine that receives input sound data representing the sound detected by the at least one microphone, and (a) the second wake word engine detects one of a plurality of command keywords supported by the second wake word engine in the input sound data, and (b) when one or more playback conditions corresponding to the detected command keyword are satisfied, is configured to generate a command keyword event. Each of the plurality of command keywords is a respective playback command. The device includes a command keyword engine that detects a first command keyword via the command keyword engine and determines whether one or more playback conditions corresponding to the first command keyword are satisfied. Based on (a) detecting the first command keyword and (b) determining that one or more playback conditions corresponding to the first command keyword are satisfied, the device generates a command keyword event corresponding to the first command keyword via the command keyword engine. In response to the command keyword event and in response to determining that one or more playback conditions are satisfied, the device executes a first playback command corresponding to the first command keyword.

[0028] A second exemplary implementation includes a network interface, one or more processors, at least one microphone configured to detect sound, at least one speaker, and a wake word engine configured to receive input sound data representing the sound detected by the at least one microphone and generate a wake word event for a voice assistant service (VAS) when the wake word engine detects a VAS wake word in the input sound data. When the VAS wake word event is generated, the wake word engine streams sound data representing the sound detected by the at least one microphone to one or more servers of the voice assistant service. The device further includes a command keyword engine configured to receive input sound data representing the sound detected by the at least one microphone. The device detects a first command keyword that is one of a plurality of command keywords supported by the device via the command keyword engine, and determines an intent based on at least one keyword via a local natural language unit (NLU). After detecting the first command keyword event and determining the intent, the device executes a first playback command corresponding to the first command keyword according to the determined intent.

[0029] Some of the embodiments described herein may refer to functions performed by a given actor such as a "user" and / or other entity, but it should be understood that this description is for illustrative purposes only. The claims should not be construed as requiring acts by any such exemplary actor unless explicitly required by the language of the claims themselves.

[0030] Furthermore, in this specification, some functions are described as being performed "based on" or "in response to" another element or function. "Based on" is to be understood as one element or function being related to another function or element. "In response to" is to be understood as one element or function being a required result of another function or element. For the sake of brevity, when a functional link exists, the function is generally described as being based on another function. However, such disclosure is to be understood as disclosing any type of functional relationship.

[0031] II. Examples of Operating Environments Figures 1A and 1B show a configuration example of a media playback system 100 (or "MPS100") in which one or more embodiments disclosed herein may be implemented. First, referring to Figure 1A, the illustrated MPS100 is associated with an exemplary home environment having a plurality of rooms and spaces, which are collectively also referred to as the "home environment", "smart home", or "environment 101". Environment 101 includes a master bathroom 101a, a master bedroom 101b (referred to herein as "Nick's room"), a second bedroom 101c, a family room or den 101d, an office 101e, a living room 101f, a dining room 101g, a kitchen 101h, and an outdoor patio 101i, and consists of a home having several rooms, spaces, and / or playback zones. In the following, specific embodiments and examples under the home environment will be described, but the techniques described herein can also be implemented in other types of environments. In some embodiments, for example, MPS100 can be implemented in one or more commercial environments (e.g., stores such as restaurants, malls, airports, hotels, retail stores), one or more vehicles (e.g., sports utility vehicles, buses, cars, ships, boats, airplanes), multiple environments (e.g., a combination of a home environment and a vehicle environment), and / or any other suitable environment where multi-zone audio is desired.

[0032] Among these rooms and spaces, the MPS100 includes one or more computing devices. Referring together to FIGS. 1A and 1B, such computing devices can include playback devices 102 (individually identified as playback devices 102a - 102o), network microphone devices 103 (individually identified as "NMD" 103a - 102i), and controller devices 104a and 104b (collectively referred to as "controller device 104"). Referring to FIG. 1B, the home environment may include additional and / or other computing devices having local network devices such as one or more smart illumination devices 108 (FIG. 1B), smart thermostats 110, and local computing device 105 (FIG. 1A). In the embodiments described below, one or more of the various playback devices 102 may be configured as portable playback devices and others may be configured as stationary playback devices. For example, headphones 102o (FIG. 1B) are portable playback devices, and the playback device 102d installed on a bookshelf may be a stationary device. As another example, the patio playback device 102c may be a battery-powered device, allowing it to be carried to various locations within environment 101 and outside environment 101 without being connected to a wall outlet or the like.

[0033] Referring to FIG. 1B, various playback devices of the MPS100, network microphones, and controller devices 102-104 and / or other network devices may be coupled to each other via a network 111 such as a LAN including a network router 109, via a point-to-point connection and / or other connections that are wired and / or wireless. For example, the playback device 102j in the den 101d (FIG. 1A) may be designated as the "left" device and may be point-to-point connected to the playback device 102a, which is also in the den 101d and may be designated as the "right" device. In related embodiments, the left playback device 102j may communicate with other network devices such as the playback device 102b, which may be designated as the "front" device, via a point-to-point connection via the network 111 and / or other connections.

[0034] As further shown in FIG. 1B, the MPS100 may be coupled to one or more remote computing devices 106 via a wide area network ("WAN") 107. In some embodiments, each remote computing device 106 may take the form of one or more cloud servers. The remote computing device 106 may be configured to interact with the computing devices in the environment 101 in various ways. For example, the remote computing device 106 may be configured to facilitate streaming and / or playback control of media content such as audio in the home environment 101.

[0035] In some implementations, various playback devices, NMDs, and / or controller devices 102-104 may be communicatively combined with at least one remote computing device related to the VAS and at least one remote computing device related to a media content service (“MCS”). For example, in the illustrated example of FIG. 1B, remote computing device 106 is associated with VAS 190 and remote computing device 106b is associated with MCS 192. In the example of FIG. 1B, only a single VAS 190 and a single MCS 192 are shown for clarity, but MPS 100 may be combined with multiple different VASs and / or MCSs. In some implementations, the VAS may be operated by one or more of AMAZON®, GOOGLE®, APPLE®, MICROSOFT®, SONOS®, or other voice assistant providers. In some implementations, the MCS may be operated by one or more of SPOTIFY®, PANDORA®, AMAZON MUSIC®, or other media content services.

[0036] As further shown in FIG. 1B, remote computing device 106 further includes a remote computing device 106c configured to perform certain operations, such as remotely facilitating media playback functions, managing device and system status information, and directing communication between the devices of MPS 100 and one or more VASs and / or MCSs. In one example, remote computing device 106c provides a cloud server for one or more SONOS Wireless HiFi System.

[0037] In various implementations, one or more of the playback devices 102 may take the form of, or include, an on-board (e.g., integrated) network microphone device. For example, the playback devices 102a - e each include, or correspond to, an NMD 103a - e. Here, a playback device equipped with an NMD is referred to as a playback device or an NMD, unless otherwise specified. In some cases, one or more of the NMDs 103 may be stand-alone devices. For example, the NMDs 103f and 103g may be stand-alone devices. In a single NMD, components and functions included in playback devices, such as speakers and related electronic devices, may be omitted. For example, in such cases, a stand-alone NMD may not perform audio output, or may perform limited audio output (e.g., relatively low-quality audio output) even if it can output audio.

[0038] The various playback devices of the MPS100 and the network microphone devices 102 and 103 may each be associated with a unique name, which may be assigned to each device by the user, such as during the setup of one or more of these devices. For example, as shown in the illustration of FIG. 1B, since the playback device 102d is physically located on the bookshelf, the user may name it "bookshelf". Similarly, since the NMD 103f is physically located on the island counter in the kitchen 101h (FIG. 1A), the name "island" may be assigned. Among the playback devices, names corresponding to zones or rooms may be assigned. For example, the playback devices 102e, 102l, 102m, and 102n may be named "bedroom", "dining room", "living room", and "office", respectively. Furthermore, a specific playback device can have a functionally descriptive name. For example, the playback devices 102a and 102b are assigned the names "right" and "front", respectively, because these two devices are configured to provide specific audio channels during media playback in the zone of the den 101d (FIG. 1A). The patio playback device 102c may be named "portable" because it is battery-powered and / or can be easily carried to different areas of the environment 101. Other naming rules are also possible.

[0039] As described above, the NMD can detect and process sounds from the environment, such as the sound of conversations of people around the NMD mixed with background noise. For example, when the NMD detects a sound in the environment, the NMD can process the detected sound to determine whether the sound contains speech that includes the NMD and ultimately voice input intended for a specific VAS. For example, the NMD can identify whether the sound contains a wake word associated with a specific VAS.

[0040] In the illustrated example of FIG. 1B, NMD103 is configured to interact with VAS190 on the network via network 111 and router 109. The interaction with VAS190 is initiated, for example, when the NMD identifies a potential wake word in the detected sound. This identification causes a wake word event to occur and initiates the transmission of the sound data detected by the NMD to VAS190. In some embodiments, various local network devices 102-105 (FIG. 1A) and / or remote computing devices 106c of MPS100 may exchange various feedback, information, instructions, and / or related data with a remote computing device associated with the selected VAS. Such information exchange may be related to a transmission message including voice input or may be independent. In one embodiment, the remote computing device(s) and MPS100 may exchange data via a communication path as described herein and / or using the metadata exchange channel described in U.S. Application No. 15 / 438,749, filed Feb. 21, 2017, entitled "Voice Control of a Media Playback System". By reference to U.S. Application No. 15 / 438,749, the entire content thereof is incorporated herein by reference.

[0041] When receiving a stream of sound data, VAS190 determines whether there is voice input in the data stream from NMD. If so, VAS190 also determines the intent of the terms included in the voice input. VAS190 then returns a response to MPS100, but this response is sent directly to the NMD that triggered the wake word event. This response is based on VAS190's determination that there is intent in the voice input. As an example, in response to receiving voice input with the command "Play Hey Jude by The Beatles", VAS190 may determine that the basic intent of the voice input is to start playback, and may further determine that the intent of the voice input is to play the specific song "Hey Jude". After these determinations, VAS190 may send a command to a specific MCS192 to obtain the content (i.e., the song "Hey Jude"), and that MCS192 may subsequently provide this content directly to MPS100 or indirectly via VAS190 (e.g., stream provision). In some embodiments, VAS190 may send a command to MPS100 such that MPS100 itself obtains the content from MCS192.

[0042] In certain embodiments, if voice input is identified in the voice detected by two or more NMDs arranged in proximity to each other, the NMDs can perform arbitration with each other. For example, the NMD-equipped playback device 102d in environment 101 (Figure 1A) is in proximity to the NMD-equipped playback device 102m in the living room, and both devices 102d, 102m may at least sometimes detect the same sound at the same time. In such a case, arbitration is necessary as to which device is responsible for sending the sound data detected by the remote VAS. An example of arbitration between NMDs is described, for example, in the previously described US Application No. 15 / 438,749.

[0043] In one embodiment, the NMD may be associated with a playback device that does not include the NMD, either by designation or by default. For example, the island NMD 103f in the kitchen 101h (FIG. 1A) may be assigned to the dining room playback device 102l that is relatively close to the island NMD 103f. In fact, in response to the remote VAS receiving an audio input from the NMD, the NMD may instruct the assigned playback device to generate audio. Here, in response to the user speaking a command to play a specific song, album, playlist, etc., an audio input is sent from the NMD to the VAS. Details regarding assigning the NMD or the playback device as a designated device or a default device are described, for example, in the previously described U.S. patent application specification.

[0044] Additional aspects related to the different components of the exemplary MPS 100, and how the different components interact to provide a media experience to the user, are described in the following sections. The discussion here generally refers to the exemplary MPS 100, but the techniques described here are not limited specifically to the applications within the home environment described above. For example, the techniques described here are also useful in other home environment configurations that include more or fewer of any of the playback devices, network microphones, and / or controller devices 102 - 104. For example, the techniques described here can be utilized within an environment having a single playback device 102 and / or a single NMD 103. In such a case, the network 111 (FIG. 1B) can be eliminated, and the single playback device 102 and / or the single NMD 103 may communicate directly with the remote computing devices 106 - d. In one embodiment, a communication network (e.g., an LTE network, a 5G network, etc.) may communicate with various playback devices, network microphones, and / or controller devices 102 - 104 independently of the LAN.

[0045] a. Examples of Playback Devices and Network Microphone Devices Figure 2A is a functional block diagram showing a particular side of the playback device 102 of the MPS100 of FIGS. 1A and 1B. As shown, the playback device 102 includes various components, each of which will be described in more detail below, and the various components of the playback device 102 are operably combined with each other via a system bus, a communication network, or some other connection mechanism. In the illustrated example of FIG. 2A, the playback device 102 includes components that support the functions of the NMD and may thus be referred to as an "NMD-equipped" playback device, such as an example of the NMD103 shown in FIG. 1A.

[0046] As shown, the playback device 102 includes at least one processor 212, which may be a clock-driven computing component configured to process input data according to instructions stored in a memory 213. The memory 213 is configured to store instructions executable by the processor 212 and is a tangible, non-transitory, computer-readable medium. For example, the memory 213 is a data storage capable of loading software code 214 executable by the processor 212 to implement a particular function.

[0047] In one example, these functions include the function of the playback device 102 (which may be another playback device) to acquire audio data from an audio source. In another example, the function includes the playback device 102 transmitting audio data, detected sound data (e.g., corresponding to an audio input), and / or other information to another device on the network via at least one network interface 224. In yet another example, the function may include the playback device 102 causing one or more other playback devices to play audio in synchronization with the playback device 102. In yet another example, the function includes enabling the playback device 102 to pair or otherwise couple with one or more other playback devices to create a multi-channel audio environment. Although many other function examples are possible, some of them will be described below.

[0048] As described above, certain functions include the playback device 102 synchronizing the playback of audio content with one or more other playback devices. During synchronized playback, a listener cannot perceive the time difference between the playback of audio content by the synchronized playback devices. The specification of U.S. Patent No. 8,234,395, filed on April 4, 2004, is titled "System and method for synchronizing operations among a plurality of independently clocked digital data processing devices" and describes in more detail several examples related to the synchronization of audio playback among playback devices.

[0049] To facilitate the playback of audio, the playback device 102 includes an audio processing component 216 configured to process the audio before the playback device 102 renders the audio. For this purpose, the audio processing component 216 includes one or more digital-to-analog converters ("DACs"), one or more audio preprocessing components, one or more audio enhancement components, one or more digital signal processors ("DSPs"), and the like. In some embodiments, one or more of the audio processing components 216 may be sub-components of the processor 212. The audio processing component 216 receives, processes, or otherwise intentionally modifies analog and / or digital audio to generate an audio signal for playback.

[0050] The generated audio signal is then sent to one or more amplifiers 217 for amplification and is played back via one or more speakers 218 operably combined with the amplifier 217. The audio amplifier 217 may include components configured to amplify the audio signal to a level for driving one or more speakers 218.

[0051] Each of the speakers 218 may include a transducer (e.g., a "driver"), or the speaker 218 as a speaker group may include a complete speaker system including an enclosure having one or more drivers. A particular driver of the speaker 218 may include, for example, a subwoofer (e.g., for low frequencies), a midrange driver (e.g., for mid frequencies), and / or a tweeter (e.g., for high frequencies). In some cases, the transducer may be driven by a respective corresponding audio amplifier of the audio amplifier group 217. In some embodiments, the playback device does not include the speaker 218 and instead may include a speaker interface for connecting the playback device to an external speaker. In a particular embodiment, the playback device does not include either the speaker 218 or the audio amplifier 217 and instead may include an audio interface (not shown) for connecting the playback device to an external audio amplifier or an audio-visual receiver.

[0052] In addition to generating an audio signal for playback by the playback device 102, the audio processing component 216 may be configured to process audio transmitted to one or more other playback devices via the network interface 224 for playback. In an exemplary scenario, the audio content processed and / or played back by the playback device 102 may be received from an external source via the audio line-in interface (e.g., an automatically detected 3.5 mm audio line-in connection) of the playback device 102 (not shown) or via the network interface 224 as described below.

[0053] As shown, at least one network interface 224 can take the form of one or more wireless interfaces 225 and / or one or more wired interfaces 226. The wireless interface may provide a network interface function for the playback device 102 to wirelessly communicate with other devices (e.g., other playback devices (if any), NMDs (if any), and / or controller devices (if any)) according to a communication protocol (e.g., any wireless standard including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G mobile communication standards, etc.). The wired interface may provide a network interface function for the playback device 102 to communicate with other devices in a wired connection according to a communication protocol (e.g., IEEE 802.3). The network interface 224 shown in FIG. 2A includes both wired and wireless interfaces, but the playback device 102 may include only a wireless interface or only a wired interface in some embodiments.

[0054] Generally, the network interface 224 facilitates the data flow between the playback device 102 and one or more other devices on the data network. For example, the playback device 102 may be configured to receive audio content via the data network from one or more other playback devices, network devices within a LAN, and / or an audio content source via a WAN such as the Internet. In one example, the audio content and other signals transmitted and received by the playback device 102 may be transmitted in the form of digital packet data consisting of an Internet Protocol (IP)-based source address and an IP-based destination address. In such a case, the network interface 224 may be configured to analyze the digital packet data so that the data directed to the playback device 102 is properly received and processed by the playback device 102.

[0055] As shown in FIG. 2A, the playback device 102 also includes an audio processing component 220 that is operably combined with one or more microphones 222. The microphones 222 are configured to detect sound (i.e., acoustic waves, also referred to as sound) in the environment of the playback device 102, and that sound is provided to the audio processing component 220. More specifically, each microphone 222 is configured to detect sound and convert the detected sound into a digital signal or an analog signal, and further, based on the detected sound, cause the audio processing component 220 to perform various functions, as will be described in more detail below. In one embodiment, the microphones 222 are arranged as an array of multiple microphones (e.g., an array of six microphones). Also, in one embodiment, the playback device 102 includes more than six microphones (e.g., eight microphones or twelve microphones) or six or fewer microphones (e.g., four microphones, two microphones, or a single microphone).

[0056] In operation, the audio processing component 220 is generally configured to detect and process sounds received via the microphone 222, identify potential voice inputs among the detected sounds, and extract the detected sound data, thereby enabling the processing of voice inputs identified among the sound data detected by a VAS such as VAS 190 (FIG. 1B). The audio processing component 220 includes various components, for example, one or more analog-to-digital converters, an acoustic echo canceller ("AEC"), a spatial processor (e.g., one or more multi-channel Wiener filters, one or more other filters, and / or one or more beamformer components), one or more buffers (e.g., one or more circular buffers), one or more wake word engines, one or more voice extractors, and / or one or more voice processing components (e.g., components capable of recognizing the voices of a specific user or multiple specific users in a household). In an exemplary embodiment, the audio processing component 220 includes one or more DSPs or modules for one or more DSPs. In this regard, a particular audio processing component 220 can also have specific parameters (e.g., gain and / or spectral parameters) that are modified or otherwise adjusted to implement a particular function. In some embodiments, one or more of the audio processing components 220 may be sub-components of the processor 212.

[0057] As further shown in FIG. 2A, the playback device 102 also includes a power component 227. The power component 227 includes at least an external power interface 228 and may be combined with a power source (not shown) via a power cable or the like that physically connects the playback device 102 to an outlet or other external power source. Other power components include, for example, a transformer, a converter, etc., for setting the power.

[0058] In some embodiments, the power component 227 of the playback device 102 may further include an internal power source 229 (e.g., one or more batteries) configured to supply power to the playback device 102 without a physical connection to an external power source. When the internal power source 229 is provided, the playback device 102 can operate without relying on an external power source. In some such embodiments, the external power source interface 228 may be configured to facilitate charging of the internal power source 229. As described above, a playback device with an internal power source may be referred to herein as a "portable playback device." On the other hand, a playback device that operates using an external power source is referred to as a "stationary playback device," although it can actually be moved within a home or the like.

[0059] The playback device 102 may further include a user interface 240, thereby facilitating user interaction, and may further be coordinated with user interaction facilitated by one or more controller devices 104. In various embodiments, the user interface 240 may include one or more physical buttons, or may support a graphical interface that provides a touch-sensitive screen(s) and / or surface(s) enabling direct user input. The user interface 240 may further include one or more of lights (e.g., LEDs) and speakers that provide visual and / or audio feedback.

[0060] As an example, FIG. 2B shows the housing 230 of the playback device 102, and includes a user interface in the form of a control area 232 in the upper portion 234 of the housing 230. The control area 232 includes buttons 236a - c for controlling audio playback, volume level, etc. The control area 232 is also provided with a button 236d for switching the microphone 222 between an on state and an off state.

[0061] As further shown in FIG. 2B, the control area 232 is at least partially surrounded by an opening formed in the upper surface portion 234 of the housing 230, through which a microphone 222 (not visible in FIG. 2B) receives sound in the environment of the playback device 102. The microphone 222 may be disposed along and / or at various positions on or in the upper surface portion 234 or other areas of the housing 230 so as to detect sound from one or more directions with respect to the playback device 102.

[0062] By way of example, Sonos, Inc. sells certain playback devices that can implement the specific embodiments disclosed herein, including "PLAY:1", "PLAY:3", "PLAY:5", "PLAYBAR", "CONNECT:AMP", "PLAYBASE", "BEAM", "CONNECT", and "SUB". Other playback devices issued in the past, present, and / or future may be used additionally or alternatively to implement the playback devices of the exemplary embodiments disclosed herein. Further, the playback device is not limited to the examples shown in FIGS. 2A or 2B or to Sonos products. For example, the playback device may include or take the form of a wired or wireless headset and may operate as part of the MPS 100 via a network interface or the like. As another example, the playback device may include or be able to interact with a docking station for a personal mobile media playback device. In yet another example, the playback device can be integrated with other devices and components used indoors and outdoors, such as a television or lighting fixture.

[0063] Figure 2C is a diagram of an exemplary audio input 280 that can be processed by an NMD or an NMD-equipped playback device. The audio input 280 can include a keyword portion 280a and an utterance portion 280b. The keyword portion 280a can include a wake word or a command keyword. In the case of a wake word, the keyword portion 280a corresponds to the detected sound that triggers the wake word. The utterance portion 280b corresponds to the detected sound that potentially includes the user's request following the keyword portion 280a. The utterance portion 280b can be processed to identify the presence of any word in the sound data detected by the NMD in response to an event triggered by the keyword portion 280a. In various implementations, the underlying intent can be determined based on the words in the utterance portion 280b. In certain implementations, the underlying intent can also be, or at least partially based on, specific words within the keyword portion 280a, such as when the keyword portion includes a command keyword. In either case, the words can correspond to one or more commands, as well as specific commands and specific keywords. The keyword in the audio utterance portion 280b can be, for example, a word that identifies a specific device or group in the MPS100. For example, in the illustrated example, the keyword in the audio utterance portion 280b can be one or more words that identify one or more zones where music is to be played, such as the living room and the dining room (FIG. 1A). In some cases, the utterance portion 280b can include additional information, such as a detected pause or silence (e.g., a period of not speaking) between the words spoken by the user, as shown in FIG. 2C. The pause can define the position of a separate command, keyword, or other information spoken by the user within the utterance portion 280b.

[0064] Based on the criteria of specific commands, NMD and / or the remote VAS can act as a result of identifying one or more commands of the voice input. The command criteria may be based, inter alia, on including specific keywords in the voice input. In addition or alternatively, the command criteria of the commands may include the identification of one or more control state variables and / or zone state variables in conjunction with the identification of one or more specific commands. The control state variables can include, for example, indicators identifying the level of volume, queues associated with one or more devices, and playback states such as whether the device is playing the queue or paused. The zone state variables can include, for example, indicators identifying which zone players are grouped, if any.

[0065] In some implementations, the MPS100 is configured to temporarily reduce the volume of the audio content being played when detecting a specific keyword such as a wake word in the keyword portion 280a. The MPS100 can restore the volume after processing the voice input 280. Such a process can be called ducking, an example of which is disclosed in U.S. Patent Application Publication No. 15 / 438,749, which is hereby incorporated by reference in its entirety.

[0066] Figure 2D shows an exemplary sound sample. In this example, the sound sample corresponds to a sound data stream (e.g., one or more audio frames) associated with a spotted wake word or command keyword within the keyword portion 280a of FIG. 2A. As shown, the exemplary sound sample includes sounds detected in the NMD environment that can be referred to as (i) a pre-roll portion (before the event) (between times t0 and t1), immediately before the wake word or command word is spoken, (ii) a wake meter portion (between times t1 and t2), while the wake word or command word is being uttered, and / or (iii) a post-roll portion (after the event) (between times t2 and t3), after the wake word or command word has been spoken. Other sound samples are possible. In various implementations, the manner of the sound sample can be evaluated according to an acoustic model aimed at mapping mel / spectral features to the phonemes of a given language model for further processing. For example, automatic speech recognition (ASR) can include such a mapping for command-keyword detection. In contrast, the wake word detection engine can be precisely tuned to identify a particular wake word and a downstream action that invokes the VAS (e.g., by targeting only non-words in the audio input processed by the playback device).

[0067] The ASR for command keyword detection can be adjusted to accommodate a wide range of keywords (e.g., 5, 10, 100, 1,000, 10,000 keywords). Command keyword detection, in contrast to wake word detection, can include supplying the ASR output to an on-board local NLU that determines with the ASR when a command word event occurred. In some implementations described later, the local NLU can determine the intent based on one or more other keywords in the ASR output generated by a particular audio input. In these or other implementations, the playback device can act on the detected command keyword event only when the playback device determines that certain conditions, such as environmental conditions (e.g., low background noise), are met.

[0068] b. Configuration Example of the Reproduction Device Figures 3A to 3E show exemplary configurations of the reproduction device. First, referring to Figure 3A, in some exemplary embodiments, a single reproduction device may belong to a zone. For example, the patio reproduction device 102c (Figure 1A) may belong to Zone A. In some embodiments described below, a plurality of reproduction devices can be "bonded" to form a "bonded pair", and they can together form one zone. For example, the reproduction device 102f (Figure 1A) named "Bed 1" in Figure 3A and the reproduction device 102g (Figure 1A) named "Bed 2" in Figure 3A may be bonded to form Zone B. Each of the bonded reproduction devices has different reproduction responsibilities (e.g., channel responsibilities). In another embodiment described later, a plurality of reproduction devices can be integrated to form one zone. The integrated reproduction devices 102d, 102m may not particularly have different reproduction responsibilities assigned thereto. That is, the integrated reproduction devices 102d, 102m can of course reproduce the audio content synchronously, but they may also reproduce the audio content in the same way as when they are not integrated respectively.

[0069] For control, each zone of the MPS 100 may be represented as a single user interface ("UI") entity. For example, as displayed by the controller device 104, Zone A may be provided as a single entity named "Portable", Zone B may be provided as a single entity named "Stereo", and Zone C may be provided as a single entity named "Living Room".

[0070] In various embodiments, a zone may inherit the name of one of the playback devices as the space to which the zone belongs. For example, Zone C may inherit the living room as the name of playback device 102m (as shown in the figure). In another example, Zone C may instead claim the bookshelf as the name of playback device 102d. In a further example, Zone C can take a name that combines in some form the playback device 102d on the bookshelf and the playback device 102m in the living room. The name selected can be chosen by the user via input at the controller device 104. In some embodiments, a zone may be given a name different from the playback devices belonging to that zone. For example, Zone B in FIG. 3A is named "Stereo", but there is no playback device with this name in Zone B. In one example, Zone B is a single UI entity representing a single device named "Stereo" composed of the constituent devices "Bed 1" and "Bed 2". In one embodiment, the playback device for Bed 1 may be playback device 102f in the master bedroom 101h (FIG. 1A), and the playback device for Bed 2 may also be playback device 102g in the same master bedroom 101h (FIG. 1A).

[0071] As described above, combined playback devices may have different playback responsibilities, such as responsibility for playing a particular audio channel. For example, as shown in FIG. 3B, devices 102f and 102g for Bed 1 and Bed 2 may be combined to produce or enhance the stereo effect of audio content. In this example, the playback device 102f for Bed 1 may be configured to play the left-channel audio components, and the playback device 102g for Bed 2 may be configured to play the right-channel audio components. In some embodiments, such a stereo combination is also referred to as "pairing".

[0072] Furthermore, the playback devices configured to be coupled can have additional and / or different respective speaker drivers. As shown in FIG. 3C, the playback device 102b named "front" may be coupled with the playback device 102k named "sub". Note that the "front" playback device 102b may render the mid - to - high frequency range, and the "sub" playback device 102k may render the low frequency range, such as a subwoofer. When the coupling is released, the "front" playback device 102b may be configured to render the full - range frequency. As another example, in FIG. 3D, the "front" and "sub" playback devices 102b and 102k are shown further coupled with the right and left playback devices 102a and 102j respectively. In some embodiments, the right and left playback devices 102a and 102j may form the surround or "satellite" channels of a home theater system. The coupled playback devices 102a, 102b, 102j, 102k may form a single zone D (FIG. 3A).

[0073] In some embodiments, the playback devices may also be "merged". Different from the coupled playback devices, the merged playback devices have no assigned playback responsibility and render the full range of audio content within the possible range of each playback device. Nevertheless, multiple merged playback devices may be provided as a single UI entity (i.e., a zone as described above). For example, in FIG. 3E, the living - room playback devices 102d and 102m are merged, and these playback devices will be provided as a single UI entity of zone C. In one embodiment, the playback devices 102d and 102m may play audio synchronously, while outputting the full range of audio content within the range that each playback device 102d and 102m can render.

[0074] In some embodiments, a stand-alone NMD may itself join a zone. For example, the NMD 103h in FIG. 1A is named "closet" and forms Zone I in FIG. 3A. Also, the NMD can be combined or merged with other devices to form a zone. For example, the NMD device 103f named "island" is combined with the playback device 102i kitchen, and together they may be named "kitchen" to form Zone F. Details regarding assigning an NMD or a playback device as a designated device or a default device are described, for example, in U.S. Patent Application No. 15 / 438,749, previously described. In some embodiments, a stand-alone NMD may not be assigned to a zone.

[0075] A plurality of playback devices included in a zone composed of individual devices, combined devices, and / or merged devices are arranged to form a set that is an aggregate of playback devices that synchronously play audio. Such a set of playback devices may be referred to as a "group", "zone group", "sync group", or "playback group". In response to an input provided via the controller device 104, the plurality of playback devices are dynamically grouped (grouping) and ungrouped (ungrouping) to form new or different groups that synchronously play audio content. For example, referring to FIG. 3A, Zone A can be grouped with Zone B to form a zone group that includes the playback devices of the two zones. As another example, Zone A may be grouped with one or more other zones C - I. Zones A - I can be grouped and ungrouped in a number of ways. For example, among Zones A - I, 3, 4, 5, or more (e.g., all) of the zones may be grouped. When grouped, the individual playback devices or combined playback devices in a zone can play audio synchronously with each other as described in the previously mentioned U.S. Patent No. 8,234,395. The grouped playback devices or combined playback devices are an example of an association between a portable playback device and a stationary playback device, and such an association is triggered in response to a trigger event as described above and will be explained in more detail below.

[0076] In various embodiments, a particular name may be assigned to a zone within the environment, and the name may be the default name of the zone within the zone group, or it may be a combination of the names of the zones within the zone group, such as "Dining Room + Kitchen" as shown in FIG. 3A. In one embodiment, the zone group may be given a unique name selected by the user, such as "Nick's Room" as also shown in FIG. 3A. The name "Nick's Room" is the name selected by the user, replacing the original room name "Master Bedroom", which was the previous name for the zone group.

[0077] In FIG. 2A, certain data may be stored in the memory 213 as one or more state variables. The variables are updated periodically and are used to describe the state of the playback zone, the playback device(s), and / or the zone group associated therewith. Also, the memory 213 may include data related to the states of other devices of the MPS 100. Such related data may be shared among the devices as needed so that one or more devices have the latest data related to the system.

[0078] In some embodiments, the memory 213 of the playback device 102 may store instances of various variable types associated with states (time-varying states). The instances of the variables can be stored with identifiers (such as tags) corresponding to the types. For example, as specific identifiers, there may be a first type "a1" for identifying a playback device in a zone, a second type "b1" for identifying a playback device in a combined state within a zone, and a third type "c1" for identifying the zone group to which the zone belongs. As a related example, in FIG. 1A, the identifier corresponding to the device named "Patio" indicates that "Patio" is the only playback device in a specific zone and is not included in any zone group. The identifier corresponding to "Living Room" indicates that "Living Room" is not grouped with other zones and includes the combined playback devices 102a, 102b, 102j, 102k. The identifier corresponding to "Dining Room" indicates that "Dining Room" is part of the "Dining Room + Kitchen" group and that devices 103f and 102i are combined. The identifier corresponding to "Kitchen" indicates the same or similar information since "Kitchen" is part of the "Dining Room + Kitchen" zone group. Examples of other zone variables and identifiers are shown below.

[0079] In yet another example, as shown in FIG. 3A, the MPS100 may include variables or identifiers that represent associations different from zones or zone groups, such as identifiers corresponding to areas. An area may include clusters of zone groups and zones that do not belong to a zone group. For example, FIG. 3A shows a first area named "First Area" and a second area named "Second Area". The first area has zones and zone groups of "Patio", "Den", "Dining", "Kitchen", and "Bathroom". The second area has zones and zone groups of "Bathroom", "Nook Room", "Bedroom", and "Living Room". In some embodiments, "area" can be used to call out clusters of zones, clusters of zone groups that share one or more zones, or other clusters of zone groups. In this case, this area is different from zone groups that do not share zones with other zone groups. Further examples of techniques for implementing an area are described in the specifications of the following U.S. patent applications. U.S. Application No. 15 / 682,506, filed on August 21, 2017, with the invention name "Room Association Based on Name", and U.S. Patent No. 8,483,853, filed on September 11, 2007, with the invention name "Controlling and manipulating groupings in a multi-zone media system". The content of each of these applications is hereby incorporated by reference in its entirety. In some embodiments, the MPS100 may not use "area", in which case the system does not store variables related to areas.

[0080] Memory 213 may be further configured to store other data. Such data may relate to an audio source accessible by playback device 102 or a playback queue to which the playback device (or several other playback devices (plural)) may be associated. In the embodiments described below, memory 213 is configured to store a set of command data for selecting a particular VAS when processing voice input. During operation, one or more playback zones in the environment of FIG. 1A may each play different audio content. For example, it is conceivable that while one user is grilling in the "patio" zone and listening to hip-hop music played by playback device 102c, another user is preparing food in the "kitchen" zone and listening to classical music played by playback device 102i. In another example, one playback zone and another playback zone may be playing the same audio content synchronously.

[0081] For example, a user may be in the "office" zone where playback device 102n is playing the same hip-hop music that playback device 102c is playing in the "patio" zone. In such a case, playback devices 102c and 102n can play hip-hop synchronously so that the user can enjoy the audio content being played at a high volume seamlessly (or at least substantially seamlessly) while moving between different playback zones. Synchronization between playback zones can be achieved in a manner similar to the synchronization between playback devices described in U.S. Patent No. 8,234,395 described above.

[0082] As described above, the zone configuration of the MPS100 may be changed dynamically. In this way, the MPS100 may support a number of configurations. For example, when a user physically moves one or more playback devices into or out of a certain zone, the MPS100 is reconfigured to accommodate the change. For example, when the user physically moves playback device 102c from the "patio" zone to the "office" zone, both playback devices 102c and 102n will be included in the "office" zone. In some cases, the user may pair or group the moved playback device 102c with those in the "office" zone and further change the names of the playback devices in the "office" zone using, for example, one controller device 104 and / or voice input. As another example, when one or more playback devices 102 are moved to a specific space in a home environment that is not yet a playback zone, the moved playback device(s) may have its name changed or be associated with the playback zone of the specific space.

[0083] Furthermore, multiple different playback zones of the MPS100 can be dynamically combined into a zone group or split into independent playback zones. For example, the "Dining Room" zone and the "Kitchen" zone may be grouped together into a zone group for a dinner party such that playback devices 102i and 102l render audio content synchronously. As another example, the combined playback device in the "Den" zone may be split into (i) a "TV" zone and (ii) another "Listening" zone. The "TV" zone may include the "front" playback device 102b. The "Listening" zone may include the grouped, paired, or merged right, left, and sub playback devices 102a, 102j, 102k as described above. By splitting the "Den" zone in this way, one user can listen to music in the "Listening" zone, which is an area of the living room space, and another user can watch TV in another area of the living room space. In a related example, a user can use either NMD103a or 103b (FIG. 1B) to control the "Den" zone before it is separated into the "TV" zone and the "Listening" zone. Once separated, the "Listening" zone is controlled, for example, by a user in the vicinity of NMD103a, and the "TV" zone is controlled, for example, by a user in the vicinity of NMD103b. However, as described above, either of the NMD103s may be configured to control various playback devices of the MPS100 and other devices.

[0084] c. Examples of Controller Devices Figure 4 is a functional block diagram showing an example of one selected among the controller devices 104 of the MPS100 of FIG. 1A. Such a controller device is hereinafter referred to as a "control device" or a "controller". The controller device shown in FIG. 4 includes components generally similar to specific components of the network devices described above, such as a processor 412, a memory 413 storing program software 414, at least one network interface 424, and one or more microphones 422. As an example, the controller device may be a dedicated controller of the MPS100. In another example, the controller device may be a network device, such as an iPhone (registered trademark), an iPad (registered trademark), other smartphones, tablets, network devices (such as network computers such as PCs and Macs (registered trademarks)), etc., on which controller application software of a media playback system is installed.

[0085] The memory 413 of the controller device 104 may be configured to store controller application software and other data related to the user of the MPS100 and / or the system 100. The memory 413 may store instructions of software 414 executable by the processor 412 to implement specific functions, such as facilitating user access, control, and / or configuration of the MPS100. The controller device 104 is configured to communicate with other network devices via a network interface 424, which may take the form of a wireless interface as described above.

[0086] In one example, system information (e.g., state variables, etc.) may be communicated between the controller device 104 and other devices via the network interface 424. For example, the controller device 104 may receive information regarding the configuration of the playback zones and the configuration of zone groups in the MPS 100 from a playback device, an NMD, or other network devices. Similarly, the controller device 104 may transmit such system information to a playback device or other network devices via the network interface 424. In some examples, the other network device may be another controller device.

[0087] Also, the controller device 104 may communicate playback device control commands, such as volume adjustment and audio playback control, to the playback device via the network interface 424. As described above, changes to the configuration of the MPS 100 may also be performed by the user using the controller device 104. Configuration changes include adding / removing one or more playback devices to / from a zone, adding / removing one or more zones to / from a zone group, forming combined or merged players, separating one or more playback devices from a combined or merged playback device, and the like.

[0088] As shown in FIG. 4, the controller device 104 also generally includes a user interface 440 configured to facilitate user access and control of the MPS100. The user interface 440 may include a touch screen display or other physical interface configured to provide various graphical controller interfaces, such as the controller interfaces 540a and 540b shown in FIGS. 5A and 5B. Referring together to FIGS. 5A and 5B, the controller interfaces 540a and 540b include a playback control region 542, a playback zone region 543, a playback status region 544, a playback queue region 546, and a source region 548. The illustrated user interface is provided on a network device such as the controller device shown in FIG. 4 and is an example of an interface that may be accessed by a user to control a media playback system such as the MPS100. To provide similar control access to the media playback system, other user interfaces of various formats, styles, and interactive sequences may be implemented on one or more network devices.

[0089] The playback control region 542 (FIG. 5A) may include selectable icons (e.g., by way of touch or cursor use) that, when selected, cause the playback device within the selected playback zone or zone group to play or pause, fast forward, rewind, skip to next, skip to previous, start / end shuffle mode, start / end repeat mode, start / end crossfade mode, etc. The playback control region 542 may also include selectable icons that, when selected, change the equalization settings and / or the playback volume, among other possibilities.

[0090] The playback zone region 543 (FIG. 5B) may include the current status of the playback zones within the MPS100. The playback zone region 543 may also include the current status of zone groups, such as the "dining room + kitchen" zone group, as illustrated.

[0091] In some embodiments, the graphical display of the playback zones may include additional selectable icons for managing or configuring the playback zones of the MPS100, such as generating combined zones, generating zone groups, separating zone groups, and changing the names of zone groups.

[0092] For example, as shown in the figure, "Group" icons may be provided within each of the graphical frames of the playback zones. Selecting the "Group" icon within the graphical frame representing a zone causes other zones within the MPS100 to appear as options, one or more of which can be selected and grouped with that zone. The selected zone is grouped with that zone, and the playback device of that zone and the playback device of the selected zone are configured to play audio content synchronously. Similarly, a "Group" icon may be displayed within the graphical frame representing a zone group. In this case, selecting the "Group" icon causes the zones within the zone group to appear as options, and selecting one or more of them to ungroup them allows one or more zones to be removed from the zone group. Additionally, other interactions and implementations for grouping or ungrouping zones via the user interface are possible. The display of the playback zones in the playback zone area 543 (FIG. 5B) is dynamically updated when the configuration of the playback zones or zone groups is changed.

[0093] The playback status area 544 (FIG. 5A) can include a graphical representation of the audio content that is currently being played, previously played, or scheduled to be played next in the selected playback zone or zone group. The selected playback zone or zone group is visually distinguishable within the playback zone area 543 and / or the playback status area 544 on the controller interface. The graphical representation includes track title, artist name, album name, album year, track length, and / or other relevant information that is convenient for the user to know and is useful when controlling the MPS100 via the controller interface.

[0094] The playback queue area 546 may include a graphical representation of the audio content in the form of a playback queue associated with the selected playback zone or zone group. In some embodiments, each playback zone or zone group is associated with a playback queue, and the playback queue includes information corresponding to zero or more audio items for playback by the playback zone or zone group. For example, each audio item in the playback queue may include a Uniform Resource Identifier (URI), a Uniform Resource Locator (URL), or other identifier that is used by a playback device within the playback zone or zone group to search for and / or retrieve the audio item from a local audio content source or a network audio content source, which are then played by the playback device.

[0095] In one example, a playlist may be added to the playback queue, in which case information corresponding to each audio item in the playlist may be added to the playback queue. In another example, the audio items in the playback queue may be saved as a playlist. In another example, the playback queue may be empty or filled but "unused", in which case the playback zone or zone group is playing continuously streaming audio content such as Internet radio that can continue to play until stopped, rather than individual audio items with a finite playback time. In yet another example, the playback queue may include Internet radio and / or other streaming audio content items and is "in use" when the playback zone or zone group is playing those items. Other examples are possible.

[0096] When a playback zone or zone group is "grouped" or "ungrouped", the playback queue associated with the affected playback zone or zone group may be cleared or re-associated. For example, when a first playback zone including a first playback queue and a second playback zone including a second playback queue are grouped, the established new zone group may initially have an empty playback queue, or a playback queue including audio items from the first playback queue (when the second playback zone is added to the first playback zone), or a playback queue including audio items from the second playback queue (when the first playback zone is added to the second playback zone), or a related playback queue having a combination of audio items from both the first and second playback queues. Also, subsequently, when the established zone group is ungrouped, the resulting first playback zone may be re-associated with the previous first playback queue, or be made empty, or be associated with a new playback queue including audio items from the playback queue associated with the zone group established before the established zone group was ungrouped. Similarly, the resulting second playback zone may be re-associated with the previous second playback queue, or be made an empty playback queue, or be associated with a new playback queue including audio items from the playback queue associated with the zone group established before the established zone group was ungrouped. Other examples are possible.

[0097] In FIGS. 5A and 5B, the graphical representation of the audio content in the playback queue area 646 (FIG. 5A) may include the track title, artist name, track length, and / or other relevant information related to the audio content in the playback queue. In one example, the graphical representation of the audio content may have selectors for displaying additional selectable icons for managing and / or operating the playback queue and / or the audio content represented by the playback queue. For example, the displayed audio content can be selected to be deleted from the playback queue, moved to another position within the playback queue, played immediately, or played after the currently playing audio content. The playback queue associated with a playback zone or zone group may be stored in the memory of one or more playback devices within the playback zone or zone group, playback devices not belonging to the playback zone or zone group, and / or other specified devices. Playback using such a playback queue causes one or more playback devices to play the media items in the queue in sequential or random order.

[0098] The source area 548 may include a graphical representation of selectable audio content sources and / or selectable voice assistants associated with the corresponding VAS. The VAS may be selectively assigned. In some examples, multiple VASs, such as AMAZON's Alexa (registered trademark) and MICROSOFT's Cortana (registered trademark), may be activatable by the same NMD. In one embodiment, the user can exclusively assign a VAS to one or more NMDs. For example, the user may assign a first VAS to one or both of the living room NMDs 102a and 102b shown in FIG. 1A and a second VAS to the kitchen NMD 103f. Other examples are possible.

[0099] d. Examples of Audio Content Sources An audio source within the source region 548 is an audio content source from which audio content can be obtained and played by a selected playback zone or zone group. One or more playback devices within the zone or zone group are configured to obtain audio content from various available audio content sources (e.g., according to a URI or URL corresponding to the audio content) for playback. In one example, the audio content can be obtained directly by the playback device from the corresponding audio content source (e.g., via a line-in connection). In another example, the audio content is provided to a playback device on the network via one or more other playback devices or network devices. As will be described in detail below, in certain embodiments, the audio content can be provided by one or more media content services.

[0100] Examples of audio content sources include the memory of one or more playback devices in a media playback system such as the MPS100 of FIG. 1, a local music library on one or more network devices (e.g., a controller device, a network-enabled personal computer, or network attached storage (“NAS”)), a streaming audio service that provides audio content over the Internet (e.g., a cloud-based music service), or an audio source connected to the media playback system via a line-in input connection on a playback device or network device, among others.

[0101] In one embodiment, an audio content source may be added or removed from a media playback system such as the MPS100 of FIG. 1A. In one example, each time one or more audio content sources are added, removed, or updated, indexing of audio items is performed. Indexing of audio items includes scanning for identifiable audio items in all folders / directories shared on a network accessible to playback devices within the media playback system, generating, or updating, an audio content database consisting of metadata (e.g., title, artist, album, track length, etc.) and other relevant information such as the URI or URL of each identifiable audio item found. Other examples for managing and maintaining audio content sources are also contemplated.

[0102] FIG. 6 is a message flow diagram showing data exchange between devices of the MPS100. In step 650a, the MPS100 receives a display of selected media content (e.g., one or more songs, albums, playlists, Podcasts, videos, stations) via the control device 104. The selected media content can include, for example, media items stored locally on one or more devices connected to the media playback system (e.g., the audio source 105 of FIG. 1C) and / or media items stored on one or more media service servers (one or more of the remote computing devices 106 of FIG. 1B). In response to receiving the display of the selected media content, the control device 104 sends a message 651a to the playback device 102 (FIGS. 1A - 1C) to add the selected media content to the playback queue of the playback device 102.

[0103] In step 650b, the playback device 102 receives the message 651a and adds the selected media content to the playback queue for playback.

[0104] In step 650c, the control device 104 receives an input corresponding to a command to play the selected media content. In response to receiving the input corresponding to the command to play the selected media content, the control device 104 sends a message 651b to the playback device 102 to cause the playback device 102 to play the selected media content. In response to receiving the message 651b, the playback device 102 sends a message 651c to the computing device 106 requesting the selected media content. In response to receiving the message 651c, the computing device 106 sends a message 651d containing data corresponding to the requested media content (e.g., audio data, video data, URL, URI).

[0105] In step 650d, the playback device 102 receives the message 651d having data corresponding to the requested media content and plays the associated media content.

[0106] In step 650e, the playback device 102 optionally causes one or more other devices to play the selected media content. In one example, the playback device 102 is one of the combined zones of two or more players (Figure 1M). The playback device 102 can receive the selected media content and send all or part of the media content to other devices within the combined zone. In another example, the playback device 102 is a group coordinator and is configured to send and receive timing information from one or more other devices within the group. One or more other devices within the group can receive the selected media content from the computing device 106 and start playing the selected media content in response to a message from the playback device 102, whereby all devices within the group play the selected media content synchronously.

[0107] III. Exemplary Command Keyword Events Figures 7A and 7B are functional block diagrams showing aspects of NMD703a and NMD703 configured in accordance with embodiments of the present disclosure. NMD703a and NMD703b are collectively referred to as NMD703. NMD703 may be generally similar to NMD103 and may include similar components. As will be described in more detail below, NMD703a (FIG. 7A) is configured to locally process certain voice inputs without necessarily transmitting data representing the voice inputs to the voice assistant service. However, NMD703a is also configured to use the voice assistant service to process other voice inputs. NMD703b (FIG. 7B) is configured to process voice inputs using the voice assistant service, and local NLU or command keyword detection may or may not be restricted.

[0108] Referring to FIG. 7A, NMD703 includes a voice capture component (“VCC”) 760, a VAS wake word engine 770a, and a voice extractor 773. The VAS wake word engine 770a and the voice extractor 773 are operatively coupled to the VCC 760. NMD703a further includes a command keyword engine 771a operatively coupled to the VCC 760.

[0109] NMD703 further includes a microphone 720 and at least one network interface 720 as described above, and may also include other components such as an audio amplifier, a user interface, etc. that are not shown in FIG. 7A for clarity. The microphone 720 of NMD703a is configured to provide the detected sound S D from the environment of NMD703 to the VCC 760. The detected sound S D can take the form of one or more analog or digital signals. In an exemplary implementation, the detected sound S D may be composed of a plurality of signals associated with respective channels 762 supplied to the VCC 760.

[0110] Each channel 762 can correspond to a particular microphone 720. For example, an NMD having six microphones can have six corresponding channels. The detected sound S D of each channel can have a particular similarity to other channels, but can also differ in certain respects, which can be due to the position of the corresponding microphone of a given channel relative to the microphones of the other channels. For example, the detected sound S D of one or more channels can have a higher signal-to-noise ratio (the "SNR") of voice to background noise than the other channels.

[0111] As further shown in FIG. 7A, VCC 760 includes an AEC 763, a spatial processor 764, and one or more buffers 768. During operation, the AEC 763 receives the detected sound S D and filters or processes the sound in order to suppress echo and / or otherwise improve the quality of the detected sound S D . The processed sound can then be passed to the spatial processor 764.

[0112] The spatial processor 764 is typically configured to analyze the detected sound S D and identify certain characteristics such as the amplitude of the sound (e.g., decibel level), frequency spectrum, directionality, etc. In one aspect, the spatial processor 764, as described above, based on the similarities and differences of the constituent channels 762 of the detected sound S D detects the sound S from the potential user's speech DIt can assist in filtering or suppressing ambient noise. As one possibility, the spatial processor 764 can monitor a metric that differentiates speech from other sounds. Such metrics can include, for example, the energy within the speech band relative to background noise, and the entropy (a measure of spectral structure) within the speech band, which is typically lower in speech than in the most common background noises. In some implementations, the spatial processor 764 may be configured to determine the probability of the presence of speech, and an example of such a function is disclosed in U.S. Patent Application Publication No. 15 / 984,073, filed May 18, 2018, entitled "Linear Filtering for Noise-Suppressed Speech Detection", which is hereby incorporated by reference in its entirety.

[0113] During operation, one or more buffers 768, which may be part of the memory 213 (FIG. 2A) or separate therefrom, capture data corresponding to the detected sound S D More specifically, one or more buffers 768 capture the detected sound data processed by the upstream AEC 764 and the spatial processor 766.

[0114] Next, network interface 724 may provide this information to a remote server that may be associated with MPS100. In one aspect, the information stored in additional buffer 769 does not explicitly state any spoken content; rather, it implies certain unique characteristics of the detected sound itself. In a related aspect, the information may be communicated between computing devices, such as various computing devices of MPS100, without necessarily implicating privacy concerns. In fact, MPS100 can use this information to adapt and fine-tune audio processing algorithms, including sensitivity adjustments as described below. In some implementations, the additional buffer may include or can include a function similar to the look-back buffer disclosed in U.S. Patent Application Publication No. 15 / 989,715, titled "Determining and Adapting to Changes in Microphone Performance of Playback Devices", filed on May 25, 2018; U.S. Patent Application Publication No. 16 / 141,875, titled "Voice Detection Optimization Based on Selected Voice Assistant Service", filed on September 25, 2018; and U.S. Patent Application Publication No. 16 / 138,111, filed on September 21, 2018, with the invention name "Voice Detection Optimization Using Sound Metadata". These are hereby incorporated by reference in their entirety.

[0115] In any event, the detected sound data forms a digital representation (i.e., a sound data stream) S of the sound detected by microphone 720. DS In fact, the sound data stream S DS can take various forms. As one possibility, the sound data stream S DSIt can be composed of frames, each of which can contain one or more sound samples. The frames may be streamed (i.e., read) from one or more buffers 768 for further processing by downstream components such as the VAS wake word engine 770 and the voice extractor 773 of the NMD703.

[0116] In some implementations, at least one buffer 768 captures the sound data detected using a sliding window approach, and while a given amount of the last captured detected sound data (i.e., a given window) is retained in at least one buffer 768, the older detected sound data is overwritten when it is outside the window. For example, at least one buffer 768 can temporarily hold 20 frames of sound samples at a given time, discard the oldest frame after its expiration, and then capture a new frame that is added to the 19th previous frame of the sound samples.

[0117] In practice, when the sound data stream S DS is composed of frames, the frames can take various forms with various characteristics. As one possibility, the frames can take the form of audio frames having a specific resolution (e.g., 16-bit resolution) based on a sampling rate (e.g., 44,100 Hz). Additionally or alternatively, the frames can include information corresponding to a given sound sample defined by the frame, such as, among other examples, frequency response, power input level, SNR, microphone channel identification, and / or metadata indicating other information about a given sound sample. Thus, in some embodiments, the frames can include a portion of the sound (e.g., one or more samples of a given sound sample) and metadata regarding the portion of the sound. In other embodiments, the frames can include only a portion of the sound (e.g., one or more samples of a given sound sample) or metadata regarding the portion of the sound.

[0118] In either case, the components downstream of NMD703 can process the sound data stream S DS For example, the VAS wake word engine 770 can apply one or more identification algorithms to the sound data stream S DS (e.g., frames of streamed sound) to spot potential wake words of the detected sound S D This process may be referred to as automatic speech recognition. The VAS wake word engine 770a and the command keyword engine 771a apply different identification algorithms corresponding to their respective wake words and further generate different events based on the detection of the wake words of the detected sound S D

[0119] Exemplary wake word detection algorithms accept audio as input and provide an indication of whether a wake word is present in the audio. Many first - party and third - party wake word detection algorithms are known and commercially available. For example, an operator of a voice service can make an algorithm available for use on third - party devices. Alternatively, the algorithm may be trained to detect a specific wake word.

[0120] For example, when the VAS wake word engine 770a detects a potential VAS wake word, the VAS work word engine 770a provides an indication of a "VAS wake word event" (also referred to as a "VAS wake word trigger"). In the example illustrated in FIG. 7A, the VAS wake word engine 770a outputs a signal S VW indicating the occurrence of a VAS wake word event to the audio extractor 773.

[0121] ​In a multiple VAS implementation, the NMD 703 includes a VAS selector 774 (shown in dashed lines) that directs extraction by the audio extractor 773 when a given wake word is identified by a particular wake word engine (and corresponding wake word trigger), such as VAS wake word engine 770a and at least one additional VAS wake word engine 770b (shown in dashed lines), and selects a wake word trigger from the sound data stream S. DS to the appropriate VAS. In such an implementation, the NMD 703 may include multiple different VAS wake word engines and / or multiple different speech extractors, each supported by a corresponding VAS.

[0122] Similar to the discussion above, each VAS wake word engine 770 receives as input a sound data stream S from one or more buffers 768. DS and applying a recognition algorithm to cause a wake word trigger for the appropriate VAS. Thus, as one example, the VAS wake word engine 770a may be configured to recognize the wake word "Alexa" and invoke the AMAZON VAS on the NMD 703a when "Alexa" is spotted. As another example, the wake word engine 770b may be configured to recognize the wake word "Okay, Google" and invoke the Google VAS on the NMD 520 when "Okay, Google" is spotted. In a single VAS implementation, the VAS selector 774 may be omitted.

[0123] In response to a VAS wake word event (e.g., a signal S indicating a wake word event) VW In response to the DS For example, the audio extractor 773 may be configured to receive and format (e.g., packetize) a sound data stream S DSPacketize the frame into a message. The voice extractor 773 may include these messages M that can include voice input in real time or near real time V and send or stream them to the remote VAS via the network interface 724.

[0124] The VAS is configured to process the sound data stream S V contained in the message M sent from the NMD703. DS More specifically, the NMD703a is configured to identify the voice input 780 based on the sound data stream S. DS As described in connection with FIG. 2C, the voice input 780 can include a keyword part and an utterance part. The keyword part corresponds to the detected sound when a wake word event is caused, or results in a command keyword event when one or more specific conditions such as specific playback conditions are met. For example, when the voice input 780 includes the VAS wake word, the keyword part corresponds to the detected sound when the wake word engine 770a outputs the wake word event signal S VW to the voice extractor 773. In this case, the utterance part corresponds to the detected sound that potentially includes the user's request following the keyword part.

[0125] When a VAS wake word event occurs, the VAS first processes the sound data stream S DSThe keyword part within can be processed to verify the existence of the VAS wake word. In some cases, VAS can determine that the keyword part contains an incorrect wake word (for example, the word "Election" when the target VAS wake word is the word "Alexa"). In such a case, VAS can send a response to NMD703a instructing NMD703a to stop extracting sound data, whereby the voice extractor 773 stops further streaming of the detected sound data to VAS. The VAS wake word engine 770a can resume or continue monitoring the sound samples until it finds another potential VAS wake word, leading to another VAS wake word event. In some implementations, VAS does not process or receive the keyword part, but instead processes only the utterance part.

[0126] In any case, VAS processes the utterance part to identify the presence of any words in the detected sound data and determines the underlying intent from these words. The words can correspond to one or more commands as well as specific keywords. The keyword can be, for example, a word of voice input that identifies a specific device or group of the MPS100. For example, in the illustrated example, the keyword can be one or more words that identify one or more zones where music is to be played, such as the living room and the dining room (FIG. 1A).

[0127] To determine the intent of the word, the VAS typically communicates with one or more databases associated with the VAS (not shown) of the MPS100 and / or one or more databases (not shown). Such databases can store various users' data, analysis, catalogs, and other information for natural language processing and / or other processing. In some implementations, such databases can be updated for adaptive learning and feedback of neural networks based on voice input processing. In some cases, the utterance portion can include additional information such as detected pauses (e.g., periods of not speaking) between the words spoken by the user, as shown in FIG. 2C. The pause can define the position of separate commands, keywords, or other information spoken by the user within the utterance portion.

[0128] After processing the voice input, the VAS may send a response with instructions for performing one or more actions to the MPS100 based on the intent determined from the voice input. For example, based on the voice input, the VAS can, among other actions, start playback on one or more of the playback devices 102, control one or more of these playback devices 102 (e.g., increase / decrease volume, group / ungroup devices, etc.), or instruct the MPS100 to turn a specific smart device on / off. After receiving the response from the VAS, the wake word engine 770a of the NMD703 can resume or continue monitoring the sound data stream S DS1 until another potential wake word is found, as described above.

[0129] Generally, one or more identification algorithms applied by a specific VAS wake word engine, such as the VAS wake word engine 770a, are based on the detected sound stream S DSconfigured to analyze certain characteristics thereof and compare those characteristics to corresponding characteristics of one or more specific VAS wake words of a specific VAS wake word engine. For example, wake word engine 770a spots spectral characteristics of detected sound stream S DS that match spectral characteristics of one or more wake words of the engine, thereby determining that detected sound S D includes an audio input that includes a specific VAS wake word.

[0130] In some implementations, the one or more identification algorithms may be third-party identification algorithms (i.e., developed by a company other than the company that provides NMD703a). For example, an operator of a voice service (e.g., AMAZON) can make available respective algorithms (e.g., an identification algorithm corresponding to AMAZON's ALEXA) for use on third-party devices (e.g., NMD103), which are then trained to identify one or more wake words of a specific voice assistant service. Additionally or alternatively, the one or more identification algorithms may be first-party identification algorithms developed and trained to identify specific wake words that are not necessarily unique to a given voice service. Other possibilities exist.

[0131] As described above, NMD703a also includes a command keyword engine 771a in parallel with VAS wake word engine 770a. Similar to VAS wake word engine 770a, command keyword engine 771a can apply one or more identification algorithms corresponding to one or more wake words. A "command keyword event" is detected sound S DIt occurs when a specific command keyword is identified. In contrast to non-words that are typically used as VAS wake words, a command keyword functions as both a wake word and the command itself. For example, exemplary command keywords can correspond to playback commands (e.g., "play", "pause", "skip", etc.) as well as control commands ("turn on"), among other examples. Under appropriate conditions, NMD703a executes the corresponding command based on detecting any of these command keywords.

[0132] The command keyword engine 771a can employ an automatic speech recognition device 772. ASR 772 is configured to output an audio or phonemic representation, such as text corresponding to a word, as text based on the sound of the sound data stream S DS For example, ASR 772 can transcribe the spoken words represented in the sound data stream S DS into one or more strings of text that represent the voice input 780 as text. The command keyword engine 771 supplies the ASR output (labeled with S ASR ) to a local natural language unit (NLU) 779 that identifies a particular keyword as a command keyword for invoking a command keyword event, as described below.

[0133] As described above, in some exemplary implementations, NMD703a is configured to perform natural language processing, which can be executed using an on-board natural language processor, referred to herein as the natural language unit (NLU) 779. The local NLU 779 is configured to analyze the text output of the ASR 772 of the command keyword engine 771a to spot (i.e., detect or identify) the keyword of the voice input 780. In FIG. 7A, this output is the signal S ASRas shown. The local NLU 779 includes a library of keywords (i.e., words or phrases) corresponding to each command and / or parameter.

[0134] In one aspect, the library of the local NLU 779 includes command keywords. When the local NLU 779 identifies a command keyword included in signal S ASR the command keyword engine 771a generates a command keyword event and, assuming that one or more conditions corresponding to the command keyword are satisfied, executes a command corresponding to the command keyword in signal S ASR

[0135] Also, the library of the local NLU 779 may include, as keywords, those corresponding to parameters. Then, the local NLU 779 may determine the intent to be included from the matching keywords of the voice input 780. For example, when the local NLU finds a match of the keywords "David Bowie" and "kitchen" with a play command, the local NLU 779 may determine the intent to play David Bowie on the playback device 102i in the kitchen 101h. In contrast to the processing of the voice input 780 by the cloud-based VAS, the local processing of the voice input 780 by the NLU 779 on a local basis may require a relatively simple configuration. This is because the NLU 779 does not access a larger sound database or a processing unit with a relatively large processing capacity that the VAS generally accesses.

[0136] ​In some examples, the local NLU 779 can determine the intent by one or more slots corresponding to each keyword. For example, referring back to the playback of David Bowie in the kitchen example, when processing the voice input, the local NLU 779 can determine that the intent is to play music (e.g., intent = playMusic), while at the same time, by the first slot, it determines that David Bowie is the target content (e.g., slot1 = DavidBowie), and by the second slot, it determines that Kitchen 101h is the target playback device (e.g., slot2 = kitchen). Here, the intent ("play Music") is based on the command keyword, and the slots are parameters that limit the intent to specific target content and playback device.

[0137] In some examples, the command keyword engine 771a outputs a signal S indicating the occurrence of a command keyword event to the local NLU 779. CW In response to the command keyword event (e.g., in response to the signal S indicating the command keyword event) CW the local NLU 779 is configured to receive and process the signal S. ASR Specifically, the local NLU 779 scrutinizes the words within the signal S ASR to find keywords that match the keywords in the local NLU 779's library.

[0138] Local automatic speech recognition is expected to have some errors. In an example, ASR772 can generate a confidence score when transcribing spoken words into text, which indicates how close the spoken words of the voice input 780 are to the voice pattern of that word. In some implementations, the generation of a command keyword event is based on the confidence score of a given command keyword. For example, the command keyword engine 771a can generate a command keyword event if the confidence score of a given sound exceeds a given threshold (e.g., 0.5 on a scale of 0 to 1 indicates a high likelihood that the given voice is not a command keyword). Conversely, if the confidence score of a given sound is below a given threshold, the command keyword engine 771a does not generate a command keyword event.

[0139] Similarly, some errors are expected when performing keyword matching. In an example, the local NLU can generate a confidence score when determining an intent, which indicates how close the transcribed words of the signal S ASR are to the corresponding keywords in the local NLU's library. In some implementations, performing an operation according to the determined intent is based on the confidence score of the keywords that matched in the signal S ASR For example, NMD703 may perform an operation according to the determined intent when the confidence score of a given sound exceeds a given threshold (e.g., 5 on a scale of 0 to 1 indicates a higher likelihood that the given sound is not a command keyword). Conversely, if the confidence score of a given intent is below a given threshold, NMD703 does not perform an operation according to the determined intent.

[0140] As described above, in some implementations, phrases can be used as command keywords, in which case additional syllables are needed to determine a match (or non-match). For example, the phrase "Play music" has more syllables than "Play", so more sound pattern matches are required. Thus, command keywords that are phrases can generally be error-resistant wake words.

[0141] As described above, NMD703a generates a command keyword event (and also executes a command corresponding to the detected command keyword) only when certain conditions corresponding to the detected command keyword are met. These conditions are aimed at reducing the occurrence rate of misdetection command keyword events. For example, after detecting the command keyword "Skip", NMD703a generates a command keyword event (i.e., skips to the next track) only when certain playback conditions indicating that skipping should be executed are met. These playback conditions can include, for example, (i) a first state where a media item is being played, (ii) a second state where the queue is active, and (iii) a third state where a media item following the currently playing media item is included in the queue. If any of these conditions are not met, no command keyword event is generated (i.e., skipping is not executed).

[0142] NMD703a includes one or more state machines 775a to facilitate determining whether appropriate conditions are met. The state machine 775a transitions between a first state and a second state based on whether one or more conditions corresponding to the detected command keyword are met. In particular, for a given command keyword corresponding to a particular command that requires one or more specific conditions, the state machine 775a transitions to the first state when one or more specific conditions are met and transitions to the second state when at least one of the one or more specific conditions is not met.

[0143] In an exemplary implementation, the command condition is based on the state indicated by the state variable. As described above, the devices of the MPS100 can store state variables that describe the state of each device. For example, the playback device 102 can store state variables indicating the state of the playback device 102, such as the audio content currently being played (or paused), the volume level, and the network connection state. These state variables are updated (e.g., periodically or based on an event (i.e., when the state of the state variable changes)), and the state variables can be further shared among the devices of the MPS100 including the NMD703.

[0144] Similarly, the NMD703 may maintain these state variables (either by being implemented in the playback device or as a stand-alone NMD). The state machine 775a monitors the states indicated by these state variables and determines whether the states indicated by the appropriate state variables satisfy the command condition. Based on these determinations, the state machine 775a transitions between the first state and the second state as described above.

[0145] In some implementations, the command keyword engine 771 can be disabled unless certain conditions are met via the state machine. For example, the first and second states of the state machine 775a can operate as an enable / disable toggle for the command keyword engine 771a. In particular, while the state machine 775a corresponding to a particular command keyword is in the first state, the state machine 775a enables the command keyword engine 771a for the particular command keyword. Conversely, while the state machine 775a corresponding to a particular command keyword is in the second state, the state machine 775a disables the command keyword engine 771a for the particular command keyword. Thereby, the disabled command keyword engine 771a is the sound data stream S DSAbort the analysis. If at least one command condition is not met, NMD703a may suppress the generation of command keyword events when the command keyword engine 771a detects a command keyword. Suppression of generation can include gating, blocking, or preventing the output from the command keyword engine 771a from generating a command keyword event. Alternatively, suppressing generation may include stopping the supply of the sound data stream S DS to ASR772. Such suppression prevents a command corresponding to a detected command keyword from being executed when at least one command condition is not met. In such an embodiment, the command keyword engine 771a can continue to analyze the sound data stream S DS while the state machine 775a is in the first state, but the command keyword events are disabled.

[0146] Other exemplary conditions can be based on the output of a voice activity detector ("VAD") 765. The VAD 765 is configured to detect the presence (or absence) of voice activity in the sound data stream S DS . In particular, the VAD 765 can analyze with one or more voice detection algorithms in a specific time window prior to the keyword portion of the voice input 780 to determine whether voice activity was present in the environment in a frame corresponding to the pre-roll portion (before the event) of the voice input 780 (FIG. 2D).

[0147] VAD765 can utilize any suitable voice activity detection algorithm. Exemplary voice detection algorithms include determining whether a given frame contains one or more features or qualities corresponding to voice activity, and further determining whether those features or qualities are distinct from noise and have grown to a given extent (e.g., whether the value exceeds a threshold for a given frame). Some exemplary voice detection algorithms include filtering or reducing the noise of the frame before identifying features or qualities.

[0148] In some examples, VAD765 can determine whether voice activity is present in the environment based on one or more metrics. For example, VAD765 can be configured to distinguish between frames that contain voice activity and frames that do not. Frames that the VAD determines have voice activity can be caused by speech, regardless of whether it is in the near field or far field. In this and other examples, VAD765 can determine a count value for frames in the pre-event portion of the voice input 780 that indicate voice activity. If this count value exceeds a threshold percentage or number of frames, VAD765 can be configured to output a signal indicating that voice activity is present in the environment or to set a state variable to such a value. In addition to or instead of such a count value, other metrics can similarly be used.

[0149] The presence of voice activity in the environment can indicate that a voice input is being provided to NMD73. Thus, when VAD765 indicates that there is no voice activity in the environment (presumably indicated by a state variable set by VAD765), this can be configured as one of the command conditions for the command keyword. When this condition is met (i.e., VAD765 indicates that voice activity is present in the environment), the state machine 775a transitions to the first state, enabling it to execute a command based on the command keyword. Of course, other conditions in addition to the command keyword need to be met.

[0150] Furthermore, in some implementations, NMD703 can include a noise classifier 766. The noise classifier 766 is configured to determine sound metadata (such as frequency response, signal level, etc.) and identify signatures of sound metadata corresponding to various noise sources. The noise classifier 766 can include a neural network or other mathematical model configured to identify different types of noise in the detected sound data or metadata. One classification of noise can be speech (e.g., far-field voice). Another classification can be a particular type of speech such as background speech, examples of which are described in more detail with reference to FIG. 8. Background speech can be distinguished from other types of speech-like activity such as more general speech-like activity (e.g., cadence, pauses, or other characteristics) similar to the speech detected by VAD765.

[0151] For example, analyzing the sound metadata can include comparing one or more features of the sound metadata with known noise reference values or sample population data having known noise. For example, any feature of the sound metadata such as signal level, frequency response spectrum, etc., can be compared with a noise reference value or a value collected and averaged over the sample population. In some examples, analyzing the sound metadata includes projecting the frequency response spectrum into an eigenspace corresponding to the aggregated frequency response spectrum from a set of NMDs. Further, projecting the frequency response spectrum into the eigenspace can be performed as a preprocessing step to facilitate downstream classification.

[0152] In various embodiments, any number of different techniques for classifying noise using sound metadata can be used, such as machine learning using decision trees, or Bayesian classifiers, neural networks, or any other classification technique. Alternatively or additionally, various clustering techniques, such as K-means clustering, mean shift clustering, expectation maximization clustering, or any other suitable clustering technique can be used. The techniques for classifying noise can include one or more techniques disclosed in U.S. Patent Application Publication No. 16 / 227,308, filed on December 20, 2018, entitled "Optimization of Network Microphone Devices Using Noise Classification", which is hereby incorporated by reference in its entirety.

[0153] FIG. 8 shows a first plot 882a and a second plot 882b. The first plot 882a and the second plot 882b show the analyzed sound metadata regarding background audio. These signs shown in the plot are generated using principal component analysis (PCA). The data collected from various NMDs provides the overall distribution of the possible frequency response spectra. Generally, by analyzing the principal components, the changes in all field data can be represented in a Cartesian coordinate system. The contours shown in the plot of FIG. 8 reflect the eigen space. Each dot in the plot represents the value of a known noise projected into the eigen space (e.g., the single-frequency response spectrum from an NMD directed at a confirmed noise source). As seen in FIG. 8, these known noise values cluster when projected into the eigen space. In this example, the plot of FIG. 8 represents four vector analyses, and each vector corresponds to a respective feature. These features, taken collectively, are signs representing background audio.

[0154] Referring back to FIG. 7A, in some implementations, an additional buffer 769 (shown in dashed lines) is processed by the upstream AEC 763 and the spatial processor 764, and the detected sound S DInformation (e.g., metadata, etc.) can be stored. This additional buffer 769 may be referred to as the "sound metadata buffer". Such sound metadata may include, for example, the following: (1) frequency response data, (2) echo return loss enhancement measurement values, (3) voice direction measurement values, (4) arbitration statistics, and / or (5) voice spectrum data. In an exemplary implementation, the noise classifier 766 analyzes the sound metadata in the buffer 769 to classify the noise of the detected sound S D can be classified.

[0155] As described above, as one classification of sound, it may be background sound such as voice in a far - field or voice indicating a conversation that does not bother the NMD703. The noise classifier 766 can output a signal indicating the presence of background sound in the environment or set a state variable. Even if there is voice activity (i.e., speech) in the pre - event part of the voice input 780, it is not directed to the NMD703, indicating that it may be conversation voice within the environment. For example, a household member may say something like "Those kids will soon go to their play appointment" without intending to instruct the NMD703 with the command keyword "Play (playback)".

[0156] Furthermore, when the noise classifier indicates the presence of background sound in the environment, this condition can disable the command keyword engine 771a. In some implementations, the condition that there is no background sound in the environment (as perhaps indicated by a state variable set by the noise classifier 766) is configured as one of the command conditions for the command keyword. Thus, the state machine 775a does not transition to the first state when the noise classifier 766 indicates the presence of background sound in the environment.

[0157] Furthermore, the noise classifier 766 can determine whether background audio exists in the environment based on one or more metrics. For example, the noise classifier 766 can determine to indicate background audio with the count value of the frames in the pre-event portion of the audio input 780. If this count value exceeds a threshold percentage or number of frames, the noise classifier 766 can be configured to output a signal or set a state variable to indicate that background audio exists in the environment. In addition to or instead of such a count value, other metrics can also be used similarly.

[0158] In an exemplary implementation, the NMD703a can support a plurality of command keywords. To facilitate such support, the command keyword engine 771a can implement a plurality of identification algorithms corresponding to each command keyword. Alternatively, the NMD703a may implement an additional command keyword engine 771b configured to identify each command keyword. Further, the library of the local NLU779 can include a plurality of command keywords and be configured to search for text patterns corresponding to these command keywords in the signal S ASR as well.

[0159] Furthermore, command keywords may require different conditions. For example, "skip" may require the condition that a media item is being played, and "play" may require the opposite condition that a media item is not being played. Thus, the condition for "skip" may be different from the condition for "play". To facilitate each of these conditions, the NMD703a can implement a corresponding state machine 775a for each command keyword. Alternatively, the NMD703a may implement one state machine 775a that can correspond to each of the command keywords. Other examples are similarly possible.

[0160] In some exemplary implementations, the VAS wake word engine 770a generates a VAS wake word event when certain conditions are met. The NMD 703b includes a state machine 775b similar to the state machine 775a. The state machine 775b transitions between a first state and a second state based on whether one or more conditions corresponding to the VAS wake word are met.

[0161] For example, in some cases, the VAS wake word engine 770a can generate a VAS wake word event only if there is no background voice present in the environment before the VAS wake word event is detected. An indication of whether voice activity is present in the environment can be provided by the noise classifier 766. As described above, the noise classifier 766 can be configured to output a signal or set a state variable to indicate that there is far-field voice in the environment. Further, the VAS wake word engine 770a can generate a VAS wake word event only if voice activity is present in the environment. As described above, the VAD 765 can be configured to output a signal or set a state variable to indicate that voice activity is present in the environment.

[0162] Illustratively, as shown in FIG. 7B, the VAS wake word engine 770a is connected to the state machine 775b. The state machine 775b remains in the first state when one or more conditions are met, and one of the conditions can include that there is no voice activity in the environment. When the state machine 775b is in the first state, the VAS wake word engine 770a is enabled and is in a state where a VAS wake word event can be generated. If any of the one or more conditions are not met, the state machine 775b transitions to the second state, thereby disabling the VAS wake word engine 770a.

[0163] Furthermore, NMD703 may include one or more sensors that output a signal indicating whether one or more users are in proximity to NMD703. Exemplary sensors include temperature sensors, infrared sensors, imaging sensors, and / or capacitance sensors, among other sensors. NMD703 can use the output from such sensors to set one or more state variables indicating whether one or more users are in proximity to NMD703. The state machine 775b may then use the presence or absence thereof as a condition for the state machine 775b. For example, the state machine 775b can enable the VAS wake word engine and / or the command keyword engine 771a when at least one user is in proximity to NMD703.

[0164] To illustrate the operation of an exemplary state machine, FIG. 7C is a block diagram showing the operation of state machine 775 for an exemplary command keyword that requires one or more command conditions. In 777a, state machine 775 remains in the first state 778a while all command conditions are met. While state machine 775 remains in the first state 778a (i.e., all command conditions are met), NMD703a generates a command keyword event when a command keyword is detected by command keyword engine 771a.

[0165] In 777b, if any command condition is not met, state machine 775 transitions to the second state 778b. In 777c, state machine 775 remains in the second state 778b if any command condition is not met. While state machine 775 is in the second state 778b, NMD703a does not transition to the command keyword event operation even if a command keyword is detected by command keyword engine 771a.

[0166] Referring back to FIG. 7A, in some examples, one or more additional command keyword engines 771b may include a custom command keyword engine. A cloud service provider, such as a streaming audio service, can provide a custom keyword engine preconfigured with an identification algorithm configured to spot service-specific command keywords. These service-specific command keywords can include commands for custom service functions and / or custom names used when accessing the service.

[0167] For example, NMD703a can include a specific streaming audio service (e.g., Apple Music) command keyword engine 771b. This specific command keyword engine 771b can be configured to detect command keywords specific to a particular streaming audio service and generate a streaming audio service wake word event. For example, one command keyword may be "Friends Mix" corresponding to a command to play a custom playlist generated from the playback history of one or more "friends" in a particular streaming audio service.

[0168] Since the custom command keyword engine 771b is generally more complex than the VAS wake word engine 770a, there may be a relatively higher likelihood of false wake words occurring in the custom command keyword engine 771b than in the VAS wake word engine 770a. To mitigate this, the custom command keyword may require that one or more conditions be met before generating a custom command keyword event. Additionally, in some implementations, multiple conditions may be imposed as requirements for including the custom command keyword engine 771b in NMD703a to reduce the occurrence rate of false detections.

[0169] These custom command keyword conditions can include service-specific conditions. For example, command keywords corresponding to premium features or playlists may require a subscription as a condition. As another example, custom command keywords corresponding to a specific streaming audio service may require media items from that streaming audio service within the playback queue. Other conditions are possible as well.

[0170] To gate the custom command keyword engine based on custom command keyword conditions, NMD703a can have additional state machines 775a corresponding to each custom command keyword. Alternatively, NMD703a may implement a state machine 775a with respective states for each custom command keyword. Other examples are possible as well. These custom command conditions may depend on state variables maintained by devices within MPS100, or may depend on state variables or other data structures representing the state of a user account of a cloud service such as a streaming audio service.

[0171] Figures 9A and 9B show Table 985 which shows exemplary command keywords and corresponding conditions. As shown in the figure, exemplary command keywords can include families that have similar intents and require similar conditions. For example, the "next" command keyword has families of "skip" and "fast forward", each of which calls the skip command under appropriate conditions. The conditions shown in Table 985 are exemplary. Various implementations can use different conditions.

[0172] Referring back to FIG. 7A, in an exemplary embodiment, the VAS wake word engine 770a and the command keyword engine 771a can take various forms. For example, the VAS wake word engine 770a and the command keyword engine 771a may take the form of one or more modules stored in the memory of NMD703a and / or NMD703b (e.g., the memory 112b of FIG. 1F). As another example, the VAS wake word engine 770a and the command keyword engine 771a may take the form of a general-purpose processor or a dedicated processor, or a combination of these modules. In this regard, the plurality of wake word engines 770 and 771 may be part of the same component of NMD703a, or each wake word engine 770 and 771 may take the form of a component dedicated to a particular wake word engine. Other possibilities also exist.

[0173] To further reduce false detections, the command keyword engine 771a can utilize a relatively low sensitivity compared to the VAS wake word engine 770a. In practice, the wake word engine can include a changeable sensitivity level setting. The sensitivity level can define the similarity between a word identified in the detected sound stream S DS1 and one or more specific wake words of the wake word engine that are considered to match (i.e., trigger a VAS wake word or command keyword event). In other words, the sensitivity level can, for example, define how closely the spectral characteristics of the detected sound stream S DS2 must match the spectral characteristics of one or more wake words of the engine that triggers the wake word.

[0174] In this regard, the sensitivity level generally controls the number of false detections identified by the VAS wake word engine 770a and the command keyword engine 771a. For example, if the VAS wake word engine 770a is configured to identify the wake word "Alexa" with a relatively high sensitivity, the wake word engine 770a may flag the presence of the wake word "Alexa" even for spurious wake words such as "Election" or "Lexus". In contrast, if the command keyword engine 771a is configured with a relatively low sensitivity, spurious wake words such as "may" or "day" will not cause the command keyword engine 771a to flag the presence of the command keyword "Play".

[0175] In practice, the sensitivity level can take various forms. In an exemplary implementation, the sensitivity level takes the form of a confidence threshold that defines the minimum confidence (i.e., probability) level of the wake word engine that functions as a line separating whether a wake word event is triggered or not when the wake word engine analyzes the detected sound of that particular wake word. In this regard, a higher sensitivity level corresponds to a lower confidence threshold (and more false detections), and a lower sensitivity level corresponds to a higher confidence threshold (and fewer false detections). For example, lowering the confidence threshold of the wake word engine configures it to trigger a wake word event when it identifies a word that is less likely to be an actual particular wake word, while raising the confidence threshold configures the engine to trigger a wake word event when it identifies a word that is more likely to be an actual particular wake word. In the example, the sensitivity level of the command keyword engine 771a may be based on more confidence scores such as more of them, or the confidence score when spotting a command keyword, and / or the confidence score when determining an intent. Other examples of sensitivity levels are possible.

[0176] In an exemplary implementation, the sensitivity level parameters (e.g., the range of sensitivity) of a particular wake word engine can be updated, and this can be done in various ways. As one possibility, the VAS of a given wake word engine or another third-party provider can provide an update to the wake word engine that modifies one or more sensitivity level parameters for a given VAS wake word engine 770a to the NMD 703. In contrast, the sensitivity level parameters of the command keyword engine 771a may be configured by the manufacturer of the NMD 703a or by another cloud service (e.g., in the case of the custom wake word engine 771b).

[0177] In particular, in a specific example, when processing the voice input 780, if a command keyword is included, the NMD 703a stops transmitting any data (e.g., message M D ) representing the detected sound S V to the VAS. In an implementation including the local NLU 779, the NMD 703a can process the voice utterance part of the voice input 780 without transmitting it to the VAS even if the voice utterance part of the voice input 780 (in addition to the word part of the keyword) exists. Therefore, even if the voice input 780 (with a command keyword) is spoken to the NMD 703, privacy can be enhanced compared to an NMD of the type that uses the VAS to process all voice inputs.

[0178] As described above, the keywords in the library of the local NLU 779 correspond to parameters. These parameters can be defined to execute commands corresponding to the detected command keywords. When a keyword is recognized in the voice input 780, the command corresponding to the detected command keyword is executed according to the parameters corresponding to the detected keyword.

[0179] For example, an exemplary voice input 780 may be "Play music at a low volume", where "Play" is the command keyword part (corresponding to the play command), and "Play music at a low volume" may be the voice utterance part. When parsing this voice input 780, the NLU 779 may recognize that "low volume" is a keyword in a library corresponding to a parameter representing a certain (low) volume level. Thus, the NLU 779 may determine an intention to play at this low volume level. Then, when executing the play command corresponding to "Play", this command is executed according to a parameter representing a certain volume level.

[0180] In a second example, another exemplary voice input 780 may be "Play favorites in the kitchen", where again "Play" is the command keyword part (corresponding to the play command), and "Play favorites in the" is the voice utterance part. Analyzing this voice input 780, the NLU 779 may recognize that "favorites" and "kitchen" match keywords in its library. In particular, "favorites" corresponds to a first parameter representing a specific audio content (i.e., a specific playlist including the user's favorite audio tracks), while "kitchen" corresponds to a second parameter representing the target of the play command (i.e., the kitchen 101h zone). Thus, the NLU 779 may determine an intention to play this specific playlist in the kitchen 101h zone.

[0181] In the third example, a further exemplary voice input 780 may be "volume up", where "volume" is the command keyword part (corresponding to the volume adjustment command) and "up" may be the voice utterance part. When parsing this voice input 780, the NLU 779 may recognize that "up" is a keyword in a library corresponding to a parameter representing a certain volume increase (e.g., a 10-point increase on a 100-point scale). Thus, the NLU 779 may determine an intention to increase the volume. Then, when executing the volume adjustment command corresponding to "volume", this command is executed according to a parameter representing a certain volume increase.

[0182] In the example, certain command keywords are functionally linked to keywords in a subset of keywords of that keyword within the library of the local NLU 779, which can facilitate the analysis. For example, the command keyword "skip" is functionally linked to the keywords "fast forward" and "rewind", and may further be linked to their cognates. Thus, when the command keyword "skip" is detected in a voice input 780, when analyzing the voice utterance part of that voice input 780 by the local NLU 779, it may include determining whether the voice input 780 contains a keyword that matches a keyword that is functionally linked (rather than determining whether the voice input 780 contains a keyword that matches a keyword in the library of the local NLU 779). Since very few keywords are checked, this analysis is relatively faster than a complete search of the library. In contrast, non-speech VAS wake words such as "Alexa" do not provide an indication regarding the scope of the accompanying voice input.

[0183] Since command keywords alone do not provide sufficient information to execute the corresponding commands, some commands may require one or more parameters. For example, the command keyword "volume" may require a parameter to specify an increase or decrease in volume because the intention is unclear with just the utterance of "volume". As another example, the command keyword "group" may require two or more parameters to identify the target devices to be grouped.

[0184] Accordingly, in some exemplary implementations, when a given command keyword included in the voice input 780 is detected by the command keyword engine 771a, the local NLU 779 can determine whether the voice input 780 includes a keyword that matches a keyword in the library corresponding to the required parameters. If the voice input 780 includes a keyword that matches the required parameters, the NMD 703a proceeds to execute the command (corresponding to the given command keyword) according to the parameters specified by the keyword.

[0185] However, even if the voice input 780 includes a keyword that matches the parameters required for the command, the NMD 703a can prompt the user to provide the parameters. For example, in one example, the NMD 703a can play an audible query prompt such as "I heard the command, but more information is needed" or "Do you need any help?". Alternatively, the NMD 703a may send a query prompt to the user's personal device via a control application (e.g., the software component 132c of the control device 104).

[0186] In a further example, NMD703a can play a customized audible prompt for a detected command keyword. For example, after detecting a command keyword corresponding to a volume adjustment command (e.g., "volume"), the audible prompt can include more specific requests such as "Do you want to increase or decrease the volume?". As another example, in the case of a grouping command corresponding to the command keyword "group", the audible prompt may be "Which device do you want to group?". Support for such specific audible prompts can be made practical by supporting a relatively limited number of command keywords (e.g., less than 100), but other implementations can support more command keywords, although with a trade-off of requiring additional memory and processing power.

[0187] In an additional example, if the voice utterance portion does not include keywords corresponding to one or more essential parameters, NMD703a can execute the corresponding command according to one or more default parameters. For example, if the play command does not include a keyword indicating the target playback device 102 for playback, NMD703a can play by itself (e.g., NMD703a is implemented within a certain playback device 102), or play on one or more associated playback devices 102 (e.g., playback devices 102 within the same room or zone as NMD703a) by default. Further, in some examples, the user can configure the default parameters using a graphical user interface (e.g., user interface 430) or a voice user interface. For example, if the grouping command does not specify the playback devices 102 to be grouped, NMD703a may default to instructing two or more preset default playback devices 102 to form a synchronization group. The default parameters are stored in data storage (e.g., memory 112b (FIG. 1F)) and may be accessed when NMD703a determines that the keyword excludes a specific parameter. Other examples are possible as well.

[0188] In some cases, NMD703a transmits the voice input 780 to the VAS if the local NLU779 cannot process the voice input 780 (e.g., if the local NLU cannot find a match with the keywords in the library, or if the local NLU779 has a low confidence score regarding the intent). In the example, to trigger the transmission of the voice input 780, NMD703a, as described above, provides the sound data stream S to the voice extractor 773 DIt is possible to generate a bridging event to cause processing. That is, NMD703a generates a bridging event for triggering the voice extractor 773 without the VAS wake word being detected by the VAS wake word engine 770a (instead, based on the NLU 779 that cannot process the voice input 780 and the command keyword of the voice input 780).

[0189] Before sending the voice input 780 to the VAS (e.g., via message M V ), NMD703a may obtain confirmation from the user to send the voice input 780 to the VAS. For example, NMD703a can play an audible prompt to send the voice input to the VAS such as "Sorry, I didn't understand. Would you like to ask Alexa?" or configured otherwise. In another example, NMD703a can play an audible prompt using a VAS voice such as "Do you need help?" (i.e., a voice that is known to most users associated with a particular VAS). In such an example, the generation of the bridging event (and triggering the voice extractor 773) is conditional on a second positive voice input 780 from the user.

[0190] In a particular exemplary implementation, the local NLU 779 can process the signal S without necessarily generating a command keyword event directly by the command keyword engine 771a. That is, the automatic speech recognition 772 may be configured to perform automatic speech recognition on the sound data stream S ASR and the local NLU 779 performs processing for keyword matching without receiving a command keyword event. If it is found that the keyword of the voice input 780 matches the keyword corresponding to the command (optionally including one or more keywords corresponding to one or more parameters), NMD703a executes the command according to the one or more parameters. D ​

[0191] Furthermore, in such an example, the local NLU 779 will only generate the signal S if certain conditions are met. ASR In particular, in some embodiments, the local NLU 779 can process the signal S only when the state machine 775a is in the first state. ASR . The particular conditions may include a condition corresponding to an absence of background sound in the environment. An indication of whether background sound is present in the environment may come from the noise classifier 766. As described above, the noise classifier 766 may be configured to output a signal or set a state variable indicating the presence of far-field sound in the environment. Additionally, there may be another condition corresponding to voice activity in the environment. The VAD 765 may be configured to output a signal or set a state variable indicating that voice activity is present in the environment. Similarly, the incidence of false positive detections of commands using the direct processing approach may be mitigated using the conditions determined by the state machine 775a.

[0192] In some examples, the library of the local NLU 779 is partially customized to an individual user. In a first aspect, the library can be customized to a device that is in the home of the NMD (e.g., a home in the environment 101 (FIG. 1A)). For example, the library of the local NLU can include keywords that correspond to names of devices in the home, such as the zone names of the playback device 102 of the MPS 100. In a second aspect, the library can be customized to a user of the devices in the home. For example, the library of the local NLU 779 may include keywords that correspond to names or other identifiers of the user's favorite playlists, artists, albums, etc. The user can then reference these names or identifiers when directing voice input to the command keyword engine 771a and the local NLU 779.

[0193] In an exemplary implementation, NMD703a can locally collect the libraries of the local NLU779 within network 111 (Figure 1B). As described above, NMD703a can maintain or access state variables indicating the respective states of devices (e.g., playback device 104) connected to network 111. These state variables may include the names of various devices. For example, kitchen 101h may include playback device 101b to which the zone name "Kitchen" is assigned. NMD703a can include them in the library of local NLU779 by reading these names from the state variables and training local NLU779 to recognize them as keywords. A keyword entry of a given name can then be associated with the corresponding device of the associated parameter (e.g., by an identifier of the device such as a MAC address or IP address). Then, NMD703a can customize control commands using the parameters and direct the commands to a specific device.

[0194] In a further example, NMD703a can collect the library by discovering devices connected to network 111. For example, NMD703a can send discovery requests over network 111 according to a protocol configured for device discovery such as Universal Plug and Play (UPnP) or Zero Configuration Networking. Then, devices on network 111 can respond to the discovery requests and exchange data representing device names, identifiers, addresses, etc. to facilitate communication and control over network 111. NMD703a can include them in the library of local NLU779 by reading these names from the exchanged messages and training local NLU779 to recognize them as keywords.

[0195] In a further example, NMD703a can be collected into a library using the cloud. To explain, FIG. 10 is a schematic diagram of MPS100 and cloud network 902. Cloud network 902 includes cloud servers 906, separately identified as media playback system control server 906a, streaming audio service server 906b, and IoT cloud server 906c. Streaming audio service server 906b can represent cloud servers of different streaming audio services. Similarly, IoT cloud server 906c can represent cloud servers corresponding to different cloud services that support smart devices 990 of MPS100.

[0196] One or more communication links 903a, 903b, 903c (hereinafter referred to as "link 903") communicatively connect MPS100 and cloud server 906. Link 903 can include one or more wired networks and one or more wireless networks (e.g., the Internet). Further, similar to network 111 (FIG. 1B), network 911 communicatively couples link 903 and at least a portion of the devices of MPS100 (e.g., playback device 102, NMD103 and 703a, control device 104, and / or one or more of smart devices 990).

[0197] In some implementations, media playback system control server 906a is facilitated by NMD703a (representing one or more of NMD703a in MPS100 (FIG. 7A)) to populate the library of local NLU779. In the example, media playback system control server 906a may receive data representing a request to populate the library of local NLU779 from NMD703a. Based on this request, media playback system control server 906a can communicate with streaming audio service server 906b and / or IoT cloud server 906c to obtain user-specific keywords.

[0198] In some examples, the media playback system control server 906a may utilize a user account and / or user profile to obtain user-specific keywords. As described above, a user of the MPS 100 can set a user profile to define settings and other information within the MPS 100. Subsequently, the user profile can be registered with the user accounts of one or more streaming audio services to facilitate streaming audio from such services to the playback device 102 of the MPS 100.

[0199] Through the use of these registered streaming audio services, the streaming audio service server 906b can collect data indicating the user's saved or favorite playlists, artists, albums, tracks, etc., either through a usage history or user input (e.g., through user input specifying saved media items or favorites). This data can be stored in the database of the streaming audio service server 906b to facilitate providing specific features of the streaming audio service, such as custom playlists, recommendations, and similar functions, to the user. Under appropriate conditions (e.g., after receiving the user's permission), the streaming audio service server 906b can share that data with the media playback system control server 906a via the link 903b.

[0200] Accordingly, in various embodiments, the media playback system control server 906a can maintain or access data indicating a user's saved or favorite playlists, artists, albums, tracks, genres, etc. If the user has registered a user profile with multiple streaming audio services, the saved data can include saved playlists, artists, albums, tracks, etc. from two or more streaming audio services. Further, the media playback system control server 906a can develop a more complete understanding of the user's favorite playlists, artists, albums, tracks, etc. by aggregating data from multiple streaming audio services as compared to a streaming audio service that only has access to data generated by the use of its own service.

[0201] In addition, in some implementations, in addition to the data shared from the streaming audio service server 906b, the media playback system control server 906a can collect usage data from the MPS 100 via the link 903a after receiving the user's permission. This can include data indicating the user's saved media items or favorite media items based on the zone. In different rooms, different types of music may be preferred. For example, the user may prefer upbeat music in the kitchen 101h and smoother music in the office 101e to be able to concentrate.

[0202] The media playback system control server 906a can identify the names of playlists, artists, albums, tracks, etc. that a user is likely to refer to when providing a playback command to the NMD703a via voice input, using data indicating the user's saved or favorite playlists, artists, albums, tracks, etc. It can then transmit the data representing these names to the NMD703a via the link 903a and the network 904, and then add them as keywords to the library of the local NLU779. For example, the media playback system control server 906a can send a command to the NMD703a to include a specific name as a keyword in the library of the local NLU779. Alternatively, the NMD703a (or another device of the MPS100) can identify the names of playlists, artists, albums, tracks, etc. that a user is likely to refer to when providing a playback command to the NMD703a via voice input, and then include these names in the library of the local NLU779.

[0203] Such customization can result in different operations being performed for similar voice inputs when the voice input is processed by the local NLU779 compared to being processed by the VAS. For example, the first voice input "Alexa, play my favorites in the office" contains the VAS wake word ("Alexa"), so it can trigger a VAS wake word event. The second voice input "play my favorites in the office" contains a command keyword ("play"), so it can trigger a command keyword. Thus, the first voice input is sent by the NMD703a to the VAS, while the second voice input is processed by the local NLU779.

[0204] These voice inputs are nearly identical but can cause different operations. In particular, the VAS can determine a first playlist of audio tracks to add to the queue of the playback device 102f of office 101e, to the extent of its capabilities. Similarly, the local NLU 779 can recognize the keywords "favorite" and "kitchen" in the second voice input. Thereby, the NMD 703a executes a voice command of "play" with the <favorites playlist> and <kitchen 101h zone> parameters, whereby a second playlist of audio tracks is added to the queue of the playback device 102f of office 101e. However, since the second playlist of audio tracks can depict user's saved or preferred playlists, artists, albums, and tracks from multiple streaming audio services, and / or usage data collected by the media playback system control server 906a, the second playlist of audio tracks can include a more complete and / or accurate collection of the user's favorite audio tracks. In contrast, the VAS can utilize its relatively limited concept of the user's saved or preferred playlists, artists, albums, and tracks when determining the first playlist.

[0205] To explain, FIG. 11 shows Table 1100 that shows the content of the first and second playlists determined based on similar voice inputs but with different processing. In particular, the first playlist is determined by the VAS, while the second playlist is determined by the NMD703a (presumably in cooperation with the media playback system control server 906a). As shown, both playlists are intended to include the user's favorites, but the two playlists include audio content from different artists and genres. In particular, the second playlist is configured according to the use of the playback device 102f in Office 101e and the user's interaction with a plurality of streaming audio services, while the first playlist is based on the interaction of a plurality of users with the VAS. As a result, the second playlist is more adapted to the types of music that the user prefers to listen to in Office 101e (e.g., indie rock and folk), while the first playlist more representatively represents the interaction with the VAS as a whole.

[0206] A household can include multiple users. Two or more users can configure their respective user profiles using the MPS100. Each user profile can have a unique user account for one or more streaming audio services associated with the respective user profile. Further, the media playback system control server 906a can maintain or access data indicating the saved or preferred playlists, artists, albums, tracks, genres, etc. of each user, and this data can be associated with the user profile of that user.

[0207] In various examples, the names corresponding to user profiles are collected in the library of the local NLU 779. This can facilitate references to the stored or preferred playlists, artists, albums, tracks, or genres of a particular user. For example, when a voice input such as "Play Ann's favorites on the patio" is processed by the local NLU 779, the local NLU 779 may determine that "Ann" matches a stored keyword corresponding to a particular user. Then, when executing the playback command corresponding to that voice input, the NMD 703a adds the playlist of the favorite audio tracks of that particular user to the queue of the playback device 102c of the patio 101i.

[0208] In some cases, the voice input may not contain a keyword corresponding to a particular user, but multiple user profiles are configured by the MPS 100. In some cases, the NMD 703a may determine the user profile to use when executing commands using voice recognition. Alternatively, the NMD 703a may default to a particular user profile. Further, when executing a command corresponding to a voice input for which a particular user profile was not identified, the NMD 703a can use preferences from multiple user profiles. For example, the NMD 703a can determine a favorite playback list containing preferred or stored audio tracks from each user profile registered in the MPS 100.

[0209] The IoT cloud server 906c can be configured to provide cloud services to support the smart device 990. The smart device 990 can include various "smart" internet-connected devices such as lights, thermostats, cameras, security systems, appliances, etc. For example, the IoT cloud server 906c can provide a cloud service to support a smart thermostat, whereby the user can control the smart thermostat over the internet via a smartphone app or website.

[0210] Thus, in the example, the IoT cloud server 906c can maintain or access data related to the user's smart device 990, such as device name, settings, and configuration. Under appropriate conditions (e.g., after receiving the user's permission), the IoT cloud server 906c can share that data with the media playback system control server 906a and / or the NMD703a via the link 903c. For example, an IoT cloud server 906c providing a smart thermostat cloud service can provide data representing such keywords to the NMD703a, which facilitates collecting keywords corresponding to temperature in the library of the local NLU779.

[0211] Furthermore, in some cases, the IoT cloud server 906c can also provide keywords specific to the control of their corresponding smart devices 990. For example, an IoT cloud server 906c providing a cloud service that supports a smart thermostat can provide a set of keywords corresponding to voice control of the thermostat, such as "ambient temperature", "warmer", or "cooler", among other examples. Data representing such keywords may be sent from the IoT cloud server 906c to the NMD703a via the link 903 and the network 904.

[0212] As described above, some homes may include more than the NMD703a. In an exemplary implementation, two or more NMD703a can synchronize or update the libraries of their respective local NLU779. For example, the first NMD703a and the second NMD703a may share data representing the libraries of their respective local NLU779 using a network (e.g., the network 904) in some cases. Such sharing can facilitate, among other possible advantages, the ability of the NMD703a to respond similarly to voice input.

[0213] In some embodiments, one or more of the above components can operate in conjunction with the microphone 720 to detect and store a user's voice profile that can be associated with the user's account on the MPS100. In some embodiments, the voice profile can be stored as a variable stored in a set of command information or data tables and / or compared to them. The voice profile can include aspects of the user's voice tone or frequency and / or other unique aspects of the user, such as those described in the previously referenced U.S. Patent Application Publication No. 15 / 438,749.

[0214] In some embodiments, one or more of the above-described components can operate in conjunction with the microphone 720 to determine the location of the user in a home environment and / or relative to one or more of the NMD103s. Techniques for determining the user's location or proximity can include one or more techniques disclosed in the previously referenced U.S. Patent Application Publication No. 15 / 438,749, U.S. Patent No. 9,084,058, entitled "Sound Field Calibration Using Listener Localization," filed December 29, 2011, and U.S. Patent No. 8,965,033, entitled "Acoustic Optimization," filed August 31, 2012. Each of these applications is hereby incorporated by reference in its entirety.

[0215] IV. Exemplary Command Keyword Techniques FIG. 12 is a flowchart showing an exemplary method 1200 for executing a first playback command based on a command keyword event. The method 1100 can be executed by a network microphone device such as NMD103s (FIG. 1A) that can include the features of NMD703a (FIG. 7A). In some implementations, the NMD is implemented within the playback device, as shown by the playback device 102r (FIG. 1G).

[0216] Block 1202 of method 1200 includes monitoring, i.e., watching over, an input sound data stream for (i) a wake word event and (ii) a first command keyword event. For example, the VAS wake word engine 770a of NMD703a may apply an algorithm that identifies one or more wake words in the sound data stream S DS (FIG. 7A). Further, the command keyword engine 770 may monitor the sound data stream S DS for command keywords, perhaps using ASR772 and local NLU779 as described above in connection with FIG. 7A.

[0217] Block 1204 of method 1200 includes detecting a wake word event. Detecting a wake word event is detecting, by the VAS wake word engine 770a of NMD703a via the microphone 720, a first sound that includes a first voice input having a wake word. The VAS wake word engine 770a can detect such a wake word in the first voice input using an identification algorithm.

[0218] Block 1206 of method 1200 includes streaming sound data corresponding to the first voice input to one or more remote servers of a voice assistant service. For example, the sound extractor 773 may extract at least a portion (e.g., the wake word portion and / or the voice utterance portion) of the first voice input from the sound data stream S DS (FIG. 7A). Then, NMD703 may stream this extracted data to one or more remote servers of a voice assistant service via the network interface 724.

[0219] Block 1208 of method 1200 includes detecting a first command keyword event. For example, after detecting a second sound, the command keyword engine 771a of NMD703a may detect a first command keyword in the sound data stream S corresponding to the second voice input in the second sound. Other examples are similarly possible. D

[0220] Block 1210 of method 1200 includes determining whether one or more playback conditions corresponding to the first command keyword are satisfied. Determining whether one or more playback conditions corresponding to the first command keyword are satisfied can include determining the state of a state machine. For example, the state machine 775 of NMD703a may transition to a first state when one or more existing playback conditions corresponding to the first command keyword are satisfied, and may transition to a second state when at least one of the one or more existing playback conditions corresponding to the first command keyword is not satisfied (FIG. 7C). Exemplary playback conditions are shown in Table 985 (FIGS. 9A and 9B).

[0221] Block 1212 of method 1200 includes executing a first playback command corresponding to the first command keyword. For example, NMD703a may detect a first command keyword event and execute a first playback command based on determining that one or more playback conditions corresponding to the first command keyword are satisfied. In the example, executing the first playback command can include generating one or more instructions for causing a target playback device to execute the first playback command.

[0222] In an example, the target playback device 102 for executing the first playback command can be defined directly (explicitly) or indirectly (implicitly). For example, the target playback device 102 can be directly defined by a reference in the voice input 780 to the name of one or more playback devices (e.g., by referring to a zone or zone group name). Alternatively, the voice input may not include any reference to the name of one or more playback devices and instead may indirectly refer to the playback device 102 associated with the NMD703a. The playback device 102 associated with the NMD703a can include a playback device implementing the NMD103a, or a playback device configured to be associated (e.g., if the playback device 102 is in the same room or area as the NMD703a), as shown by the playback device 102d implementing the NMD703d (FIG. 1B).

[0223] In an example, executing the first playback operation can include sending one or more instructions over one or more networks. For example, the NMD703a can send instructions locally to one or more playback devices 102 over the network 903 to execute instructions such as transport control (FIG. 10), similar to the message exchange shown in FIG. 6. Further, the NMD703a can send a request to the streaming audio service 906b to stream one or more audio tracks to the target playback device 102 for playback over the link 903 (FIG. 10). Alternatively, the instructions may be provided internally (e.g., via a local bus or other interconnection system) to one or more software or hardware components (e.g., the electronics 112 of the playback device 102).

[0224] Furthermore, sending commands can include both local and cloud-based operations. For example, NMD703a can send commands to one or more local playback devices 102 via network 903 to add one or more audio tracks to a playback queue. Next, requests can be sent from the one or more playback devices 102 to a streaming audio service 906b to stream one or more audio tracks to a target playback device 102 for playback via link 903. Other examples are possible as well.

[0225] Figure 13 is a flowchart showing an exemplary method 1300 of executing a first playback command in response to one or more parameters based on a command keyword event. Similar to method 1100, method 1300 can be executed by a network microphone device such as NMD120 (FIG. 1A) that includes features similar to those of NMD703a (FIG. 7A). In some implementations, the NMD is implemented within a playback device, as shown by playback device 102r (FIG. 1G).

[0226] Block 1302 of method 1300 includes monitoring an input sound data stream for (i) a wake word event and (ii) a first command keyword event. For example, the VAS wake word engine 770a of NMD703a can apply an algorithm to identify one or more wake words in sound data stream S DS (FIG. 7A). Further, the command keyword engine 770 can monitor sound data stream S DS for command keywords, perhaps using ASR772 and local NLU779 as described above in connection with FIG. 7A.

[0227] Block 1304 of method 1300 includes detecting a wake word event. The detection of the wake word event is to detect, by the VAS wake word engine 770a of NMD703a, via the microphone 720, a first sound that includes a first voice input having a wake word. The VAS wake word engine 770a can use an identification algorithm to detect such a wake word in the first voice input.

[0228] Block 1306 of method 1300 includes streaming sound data corresponding to the first voice input to one or more remote servers of a voice assistant service. For example, the voice extractor 773 may extract at least a part (e.g., the wake word part and / or the voice utterance part) of the first voice input from the sound data stream S DS (FIG. 7A). Then, NMD703 may stream this extracted data to one or more remote servers of the voice assistant service via the network interface 724.

[0229] Block 1308 of method 1300 includes detecting a first command keyword event. For example, after detecting a second sound, the command keyword engine 771a of NMD703a can detect a first command keyword in the sound data stream S D corresponding to the second voice input in the second sound (FIG. 7A). Further, the local NLU 779 may detect that the second voice input includes at least one keyword among the keywords included in the library of the local NLU 779. For example, the local NLU 779 may determine whether the voice input includes a keyword that matches any of the keywords included in the library of the local NLU 779. The local NLU 779 is configured to analyze the signal S ASR to spot (i.e., detect or identify) the keyword of the voice input.

[0230] Block 1310 of method 1300 includes determining an intent based on at least one keyword. For example, local NLU 779 may determine an intent from one or more keywords in the second voice input. As described above, the keywords in the library of local NLU 779 correspond to parameters. The keywords of the voice input can indicate an intent such as playing specific audio content in a specific zone.

[0231] Block 1312 of method 1300 includes executing a first playback command according to the determined intent. In some examples, executing the first playback command may include generating one or more instructions for executing the command according to the determined intent, whereby the target playback device is caused to execute the first playback command customized by the parameters in the voice utterance part of the voice input. As described above in connection with block 1212 (FIG. 12), the target playback device 102 for executing the first playback command can be defined directly or indirectly. Further, executing the first playback operation may include sending one or more instructions via one or more networks, or may include providing the instructions internally (e.g., for playing a device component).

[0232] FIG. 14 is a flowchart showing an exemplary method 1400 of executing a first playback command according to one or more parameters based on a command keyword event corresponding to the command keyword. The command keyword event may be generated only when a specific condition is met. The first state is that there was no background audio in the environment when the command keyword was detected.

[0233] Similar to methods 1200 and 1300, method 1400 can be performed by a network microphone device such as NMD120 (FIG. 1A) that can include features of NMD703a (FIG. 7A). In some implementations, the NMD is implemented within a playback device, as shown by playback device 102r (FIG. 1G).

[0234] Block 1402 of method 1400 includes detecting sound via one or more microphones. For example, NMD703 can detect sound via microphone 720 (FIG. 7A). Further, NMD703 may process the detected sound using one or more components of VCC760.

[0235] Block 1404 of method 1400 determines that (i) the detected sound includes a voice input, (ii) the detected sound excludes background audio, and (iii) the voice input includes a command keyword.

[0236] For example, to determine whether the detected sound includes a voice input, a voice activity detector 765 analyzes the detected sound to determine the presence (or absence) of voice activity in sound data stream S DS (FIG. 7A). Further, to determine whether the detected sound does not include background audio, a noise classifier 766 analyzes the sound metadata corresponding to the detected sound to determine whether the sound metadata includes features corresponding to background audio.

[0237] Further, to determine whether the voice input includes a command keyword, a command keyword engine 771a can parse sound data stream S DS (FIG. 7A). In particular, ASR772 can convert sound data stream S DS into text (e.g., signal S ASR) can be transcribed thereto, and the local NLU 779 can determine that the word matching the command keyword is in the text where the transcription is made. In other examples, the command keyword engine 771a can use one or more keyword identification algorithms on the sound data stream S DS above. Similarly, other examples are possible.

[0238] Block 1406 of method 1400 includes executing a playback function corresponding to the command keyword. For example, the NMD can execute the playback function based on the determination that (i) the detected sound includes voice input, (ii) the detected sound does not include background voice, and (iii) the voice input includes the command keyword. Blocks 1212 and 1312 of FIGS. 12 and 13 respectively provide examples of executing the playback function.

[0239] V. Exemplary Embodiments FIGS. 15A, 15B, 15C, and 15D show exemplary inputs and outputs from an exemplary NMD configured according to aspects of the present disclosure.

[0240] FIG. 15A shows a first scenario in which the wake word engine of the NMD is configured to detect three command keywords ("play", "stop", "resume"). In this case, the local NLU is deactivated. In this scenario, the user speaks "play" as voice input to the NMD, thereby triggering a new recognition of one of the command keywords (e.g., the command keyword event corresponding to play).

[0241] Furthermore, the voice activity detector (VAD) and the noise classifier analyze 150 frames of the pre-event portion of the voice input. As shown, the VAD has detected voice in 140 of the 150 pre-event frames, indicating that voice input may be present in the detected sound. Further, the noise classifier has detected ambient noise in 11 frames, background voice in 127 frames, and fan noise in 12 frames. In this NMD, the noise classifier classifies the dominant noise source in each frame. This indicates the presence of background voice. As a result, the NMD determines not to trigger the "play" of the detected command keyword.

[0242] Figure 15B shows a second scenario in which the NMD wake word engine is configured to detect a command keyword ("play") as well as two cognates of that command keyword ("play something" and "play a song"). Again, the local NLU is disabled. In this second scenario, the user speaks "play something" as voice input to the NMD, thereby triggering a new recognition of one of the command keywords.

[0243] Furthermore, the voice activity detector (VAD) and the noise classifier analyze 150 frames of the pre-event portion of the voice input. As shown, the VAD has detected voice in 87 of the 150 pre-event frames, indicating that voice input may be present in the detected sound. Further, the noise classifier has detected ambient noise in 18 frames, background voice in 8 frames, and fan noise in 124 frames. This indicates the absence of background voice. Considering the above, the NMD determines to trigger the detected command keyword "play".

[0244] Figure 15C shows a third scenario in which the NMD wake word engine is configured to detect three command keywords ("play", "stop", "resume"). The local NLU is enabled. In this third scenario, the user has spoken the voice input "Play Beatles music in the kitchen" to the NMD, thereby triggering a new recognition of one of the command keywords (e.g., the command keyword event corresponding to play).

[0245] As shown, the ASR has transcribed the voice input as "Play Beatles music in the kitchen". Some error (e.g., "beet les" for "The Beatles") is expected when running the ASR. Here, the local NLU matches the keyword "beet les" with "The Beatles" in the local NLU library. In the local NLU library, this artist is set as a content parameter for the play command. Further, the local NLU also matches the keyword "kitchen" with "kitchen" in the local NLU library, which sets the kitchen zone as a target parameter for the play command. The local NLU generated a confidence score of 0.63428231948273443 for the intent determination.

[0246] Similarly here, the voice activity detector (VAD) and the noise classifier are analyzing 150 frames of the pre-event part of the voice input. As shown, the noise classifier has detected ambient noise in 142 frames, background voice in 8 frames, and fan noise in 0 frames. This indicates the absence of background voice. The VAD has detected voice in 112 of the 150 pre-event frames, indicating that voice input could be present in the detected sound. Here, the NMD determines to trigger the detected command keyword "play".

[0247] Furthermore, the voice activity detector (VAD) and noise classifier analyze 150 frames of the pre-event portion of the voice input. As shown, the VAD has detected voice in 140 of the 150 pre-event frames, indicating that voice input may be present in the detected sound. Additionally, the noise classifier detected ambient noise in 11 frames, background voice in 127 frames, and fan noise in 12 frames. This indicates the presence of background voice. As a result, the NMD has determined not to trigger the "play" of the detected command keyword.

[0248] Figure 15D shows a fourth scenario where the keyword engine of the NMD is not configured to spot any command keyword. Instead, the keyword engine runs ASR and passes the output of the ASR to the local NLU. The local NLU is enabled and configured to detect keywords corresponding to both commands and parameters. In the fourth scenario, the user is speaking the voice input "Play some music in the office" to the NMD.

[0249] As shown, the ASR has transcribed the voice input as "Play some music in the office". Here, the local NLU has matched the keyword "play" in the local NLU library corresponding to the play command with "play". Additionally, the local NLU has also matched the keyword "office" with "office" in the local NLU library that sets the office zone as the target parameter for the play command. The local NLU has generated a confidence score of 0.14620494842529297 for the keyword matching. In some examples, this low confidence score may cause the NMD not to accept the voice input (e.g., if this confidence score falls below a threshold such as 5).

[0250] Conclusion In the above description, among other things, various exemplary systems, methods, apparatuses, and articles of manufacture have been disclosed that include firmware and / or software executed on hardware. The above description is merely exemplary and should not be construed as limiting. For example, any or all aspects or components of firmware, hardware, and / or software are contemplated to be implemented in hardware only, software only, firmware only, or any combination of hardware, software, and / or firmware. Thus, these examples are not the only ways to implement such systems, methods, apparatuses, and articles of manufacture.

[0251] The description herein has been made with respect to exemplary environments, systems, procedures, steps, logical blocks, processes, and is further represented symbolically, and is made with respect to those that are directly or indirectly similar to the operation of a data processing device connected to a network. Such descriptions and representations of processes are used by those skilled in the art to most effectively convey the essence of their work to other skilled artisans. A number of specific details have been set forth in order to provide a thorough understanding of the description herein. However, it will be understood by those skilled in the art that the specific embodiments described herein can be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have been omitted from detailed description in order to avoid unnecessarily obscuring aspects of the embodiments. Accordingly, the scope of the present disclosure is defined not by the description of the above embodiments, but by the appended claims.

[0252] If any of the appended claims are read to cover a pure software and / or firmware implementation, at least one element in at least one instance is explicitly defined herein to include a tangible non-transitory medium such as a memory, DVD, CD, Blu-ray™, etc., that stores the software and / or firmware.

[0253] The present technology will be described in accordance with various aspects as described below. Various examples of aspects of the present technology will be described for convenience as numbered examples (such as 1, 2, 3, etc.). These are illustrative and do not limit the present technology. Any combination of the dependent examples can be incorporated into each independent example. Other examples can be shown in a similar manner.

[0254] Example 1: A method executed by a playback device comprising a network interface and at least one microphone configured to detect sound, the method comprising: monitoring an input sound data stream representing sound detected by at least one microphone for (i) a wake word event and (ii) a first command keyword event; detecting a wake word event, wherein detecting the wake word event comprises detecting a first sound via one or more microphones and then determining that the detected first sound comprises a first voice input comprising a wake word; streaming sound data corresponding to at least a portion of the first voice input via the network interface to one or more remote servers of a voice assistant service; detecting a first command keyword event, wherein detecting the first command keyword event comprises detecting a second sound via one or more microphones and then determining that the detected second sound comprises a second voice input comprising a first command keyword, wherein the first command keyword is one of a plurality of command keywords supported by the playback device; determining that one or more playback conditions corresponding to the first command keyword are satisfied; and in response to detecting the first command keyword event and determining that one or more playback conditions corresponding to the first command keyword are satisfied, executing a first playback command corresponding to the first command keyword.

[0255] Example 2: After detecting a first command keyword event, detecting a subsequent first command keyword event, where detecting the subsequent first command keyword event includes, after detecting a third sound via at least one microphone, determining that the third sound includes a third voice input including the first command keyword; determining that at least one of one or more playback conditions corresponding to the first command keyword is not satisfied; and further including canceling the execution of a first playback command corresponding to the first command keyword in response to determining that at least one playback condition is not satisfied, the method according to Example 1.

[0256] Example 3: Detecting a second command keyword event, where detecting the second command keyword event includes, after detecting a third sound via at least one microphone, determining that the third sound includes a third voice input including the second command keyword of the detected third sound; determining that one or more playback conditions corresponding to the second command keyword are satisfied; and further including executing a second playback command corresponding to the second command keyword in response to detecting the second command keyword event and determining that one or more playback conditions corresponding to the second command keyword are satisfied, the method according to any one of Examples 1 and 2.

[0257] Example 4: At least one of one or more playback conditions corresponding to the second command keyword is not a playback condition among one or more playback conditions corresponding to the first command keyword, the method according to Example 3.

[0258] Example 5: The first command keyword is a skip command, and one or more playback conditions corresponding to the first command keyword include (i) a first state where a media item is being played on a playback device, (ii) a second state where a queue is active on the playback device, and (iii) a third state where the queue includes media items that follow the media item being played on the playback device. Executing a first playback command corresponding to the first command keyword includes skipping forward in the queue to play a media item following the media item being played on the playback device, according to the method described in any of Examples 1 to 4.

[0259] Example 6: The first command keyword is a pause command, and one or more playback conditions corresponding to the first command keyword include the condition that audio content is being played on a playback device. Executing a first playback command corresponding to the first command keyword includes pausing the playback of audio content on the playback device, according to the method described in any of Examples 1 to 4.

[0260] Example 7: The first command keyword is a volume increase command, and one or more playback conditions corresponding to the first command keyword include the condition that audio content is being played on a playback device and a second state where the volume level of the playback device is not at the maximum volume level. Executing a first playback command corresponding to the first command keyword includes increasing the volume level on the playback device, according to the method described in any of Examples 1 to 4.

[0261] Example 8: One or more playback conditions corresponding to the first command keyword include a first state where there is no background audio in the detected first sound, according to the method described in any of Examples 1 to 8.

[0262] Example 9: A tangible, non-transitory computer-readable medium storing instructions executable by one or more processors to cause a playback device to execute the method according to any one of Examples 1 to 8.

[0263] Example 10: A playback device including a speaker, a network interface, one or more microphones configured to detect sound, one or more processors, and a tangible and intangible computer-readable medium storing instructions executable by the one or more processors to cause the playback device to execute the method according to any one of Examples 1 to 8.

[0264] Example 11: A method executed by a playback device comprising a network interface and at least one microphone configured to detect sound, the method comprising: monitoring an input sound data stream representing sound detected by at least one microphone for (i) a wake word event and (ii) a first command keyword event; detecting a wake word event, the detecting comprising determining that the detected first sound, after detecting a first sound via one or more microphones, comprises a first voice input that includes a wake word; streaming sound data corresponding to at least a portion of the first voice input to one or more remote servers of a voice assistant service via the network interface; detecting a first command keyword event, the detecting comprising determining that the detected second voice, after detecting a second voice via one or more microphones, comprises a second voice input that includes a first command keyword and at least one keyword, the first command keyword being one of a plurality of command keywords supported by the playback device, the first command keyword corresponding to a first playback command; determining an intent based on at least one keyword via a local natural language unit (NLU), the NLU including a pre-determined library of keywords including at least one keyword; and executing a first playback command according to the determined intent after (a) detecting the first command keyword event and (b) determining the intent.

[0265] Example 12: Detecting a second command keyword event, where detecting the second command keyword event includes, after detecting a third sound via at least one microphone, determining that the third sound includes a third voice input that includes the second command keyword; determining that a third voice input that includes the second command keyword does not include at least one other keyword from a library of predetermined keywords; and after determining that a third voice input that includes the second command keyword does not include at least one keyword from a library of predetermined keywords, streaming to one or more servers of the voice assistant service sound data representing at least a portion of the voice input that includes the second command keyword for processing by one or more remote servers of the voice assistant service, the method according to Example 11 further comprising.

[0266] Example 13: Playing an audio prompt that requests confirmation to invoke a voice assistant service for processing a second command keyword; and after playing the audio prompt, receiving data representing confirmation to invoke a voice assistant service for processing the second command keyword, where streaming sound data representing at least a portion of the voice input that includes the second command keyword to one or more servers of the voice assistant service is performed only after receiving data representing confirmation to invoke the voice assistant service, further comprising receiving.

[0267] Example 14: Detecting a second command keyword event, where detecting the second command keyword event includes, after detecting a third sound via at least one microphone, determining that the third sound includes a third voice input that includes the second command keyword; determining that the third voice input that includes the second command keyword does not include at least one other keyword from a library of predetermined keywords; and further including, after determining that the third voice input that includes the second command keyword does not include at least one keyword from a library of predetermined keywords, executing a first playback command according to one or more default parameters, the method according to any one of Examples 10 to 13.

[0268] Example 15: The first keyword of at least one keyword in the detected second sound represents a zone name corresponding to a first zone of a media playback system, and executing the first playback command according to the determined intent includes sending one or more instructions for executing the first playback command in the first zone, the media playback system including a playback device, the method according to any one of Examples 10 to 14.

[0269] Example 16: Populating a library of predetermined keywords having zone names corresponding to respective zones within a media playback system, further including each zone comprising one or more respective playback devices, the library of predetermined keywords having the zone name corresponding to the first zone of the media playback system populated, the method according to any one of Examples 10 to 15.

[0270] Example 17: Discovering a smart home device connected to a local area network via a network interface, and further including populating a library of predetermined keywords with names corresponding to each smart home device discovered on the local area network, the method according to any one of Examples 10 to 16.

[0271] Example 18: The method according to any one of Examples 10 to 17, wherein the media playback system comprises a playback device, the media playback system is registered in one or more user profiles, and the function further comprises populating a library of predetermined keywords having names corresponding to playlists designated as favorites by one or more user profiles.

[0272] Example 19: The first user profile among the one or more user profiles is associated with a user account of a first streaming audio service and a user account of a second streaming audio service, and the playlist comprises a first playlist of the first streaming audio service designated as a favorite by the user account of the first streaming audio service and a second playlist comprising audio tracks from the first streaming audio service and the second streaming audio service. The method according to Example 18.

[0273] Example 20: The method according to any one of Examples 10 to 16, wherein detecting a first command keyword further comprises determining that one or more playback conditions corresponding to the first command keyword are satisfied.

[0274] Example 21: The method according to Example 20, wherein the one or more playback conditions corresponding to the first command keyword include a first state that there is no background audio in the detected first sound.

[0275] Example 22: A tangible non-transitory computer-readable medium storing instructions executable by one or more processors to cause a playback device to perform the method according to any one of Examples 10 to 21.

[0276] Example 23: A playback device, comprising a speaker, a network interface, one or more microphones configured to detect sound, one or more processors, and a tangible or intangible computer-readable medium storing instructions executable by the one or more processors to cause the playback device to execute the method according to any one of Examples 10 to 21.

[0277] Example 24: A method executed by a playback device comprising a network interface and at least one microphone configured to detect sound, the method comprising: detecting sound via one or more microphones; determining that (i) the detected sound includes voice input, (ii) the detected sound does not include background noise, and (iii) the voice input includes a command keyword; and in response to determining that (i) the detected sound includes voice input, (ii) the detected sound excludes background noise, and (iii) the voice input includes a command keyword, executing a playback function corresponding to the command keyword.

[0278] Example 25: The detected sound is a first detected sound, and the method further comprises: detecting a second sound via at least one microphone; determining that the detected second sound includes a wake word; and after determining that the detected second sound includes a wake word, streaming voice input in the detected second sound to one or more remote servers of a voice assistant service via the network interface of the playback device.

[0279] Example 26: Determining that there is no background voice in the detected second sound, determining sound metadata corresponding to the detected sound; and analyzing the sound metadata to classify the detected sound according to one or more specific signs selected from a plurality of signs, each sign of the plurality of signs being associated with a noise source, and at least one of the signs of the plurality of signs being a background voice sign indicating background voice, the method according to any one of Example 24 and Example 25.

[0280] Example 27: Analyzing sound metadata includes classifying a frame related to the detected sound as having a specific voice signature other than the background voice signature; and comparing, if there is a frame classified with the background voice signature, the number of frames classified with a signature other than the background voice signature, the method according to Example 26.

[0281] Example 28: Determining that there is voice input in the detected sound includes detecting the voice activity of the detected sound, the method according to any one of Example 24 and Example 25.

[0282] Example 29: Detecting the voice activity in the detected sound includes determining the number of first frames associated with the detected sound as including voice; and comparing the number of first frames with the number of second frames that (a) are related to the detected sound and (b) do not indicate voice, the method according to Example 28.

[0283] Example 30: The first frame includes one or more frames generated in response to near - distance voice activity and one or more frames generated in response to far - distance voice activity, the method according to Example 29.

[0284] Example 31: A tangible non - transient computer - readable medium storing instructions executable by one or more processors to cause a playback device to execute the method according to any one of Examples 24 to 30.

[0285] Example 32: A playback device, comprising one or more microphones configured to detect sound, one or more processors, and a tangible or intangible computer-readable medium storing instructions executable by the one or more processors to cause the playback device to perform the method according to any one of Examples 24 to 30.

Claims

1. A method performed by a playback device comprising at least one microphone configured to detect sound, and first and second wake word engines configured to receive input sound data representing the sound detected by the at least one microphone, the method comprising: When the first wake word engine detects a voice assistant service (VAS) wake word in the input sound data, causing the playback device to stream sound data representing the sound detected by the at least one microphone towards one or more servers of the VAS, configured to generate a VAS wake word event, the method comprising: Detecting, by the second wake word engine, a first command word among a plurality of command words in the sound data, wherein each of the plurality of command words corresponds to various playback commands, and wherein the sound data is a first voice input, the first voice input including the first command word and a first voice utterance; Determining, via a local natural language unit (NLU), whether the first voice utterance includes at least one keyword; and Sending the voice utterance to one or more servers of a voice assistant service if the local natural language unit (NLU) cannot find a match with one or more keywords corresponding to the command word in a predetermined keyword library.

2. The method of claim 1, wherein the input sound data does not include a wake word corresponding to a remote voice assistant service (VAS).

3. Further, When it is determined that one or more playback conditions corresponding to the detected first command word are satisfied, the second wake word engine generates a command word event corresponding to the second command word detected in the second input sound data, and executing a first playback command corresponding to the second command word in response to the command word event, the method of claim 1. **Claim 4** The playback device further includes a state machine configured to transition to a first state when one or more playback conditions corresponding to the first command word are satisfied, and to transition to a second state when at least one of the one or more playback conditions corresponding to the first command word is not satisfied. Determining that one or more playback conditions corresponding to the first command word are satisfied includes determining that the state machine is in the first state, the method according to any one of claims 1 to 3. **Claim 5** The playback device further includes an additional state machine corresponding to each command word of the plurality of command words, and each additional state machine transitions to the first state when one or more playback conditions corresponding to the respective command word are satisfied, and transitions to the second state when at least one of the one or more playback conditions corresponding to the respective command word is not satisfied, the method according to claim 4. **Claim 6** Furthermore, When it is determined that at least one of the one or more playback conditions corresponding to the detected first command word is not satisfied, suppressing generation of a command word event corresponding to the detected first command word, the method according to any one of claims 1 to 5. **Claim 7** Furthermore, storing sound data representing the sound detected by the at least one microphone in a buffer, When the first wake word engine detects a first VAS wake word, the first wake word engine generates a VAS wake word event corresponding to the detected first VAS wake word, and responding to the command word event corresponding to the first command word, via a network interface, to one or more servers of the voice assistant service, a part of the buffered sound data, (i) the buffered sound data for a predetermined time preceding the first VAS wake word, and (ii) the buffered sound data representing the voice following the first VAS wake word, streaming, the method according to any one of claims 1 to 6. **Claim 8** The playback device further comprises a third wake word engine configured to receive input sound data representing the sound detected by the at least one microphone, The method is Furthermore the third wake word engine detects a specific streaming audio service wake word among a plurality of audio service wake words supported by the third wake word engine, where the plurality of audio service wake words correspond to respective streaming audio service commands, when it is determined that one or more streaming audio service conditions corresponding to the specific streaming audio service wake word are satisfied, the third wake word engine generates a streaming audio service wake word event corresponding to the specific streaming audio service wake word, and The method according to any one of claims 1 to 7, comprising executing a specific streaming audio service command corresponding to the specific streaming audio service wake word in response to the streaming audio service wake word event.

9. Furthermore, when the local natural language unit (NLU) determines that the first voice utterance includes one or more specific keywords from a predetermined keyword library, the method according to any one of claims 1 to 8, comprising executing a first playback command according to one or more parameters corresponding to the one or more specific keywords in the first voice utterance.

10. The first keyword of the one or more specific keywords represents a zone name corresponding to a first zone of the media playback system, Furthermore, the method according to any one of claims 1 to 9, comprising executing the first playback command according to the one or more parameters by sending one or more instructions for executing the first playback command in the first zone.

11. Furthermore, populating a library of predetermined keywords with zone names corresponding to respective zones within the media playback system, where each zone comprises one or more respective playback devices, the method according to any one of claims 1 to 10, wherein the library of predetermined keywords is populated with a zone name corresponding to a first zone of the media playback system.

12. Furthermore, detecting smart home devices connected to a local area network via a network interface, and populating the library of predetermined keywords with names corresponding to each of the smart home devices detected on the local area network The method according to any one of claims 1 to 11, comprising **Claim 13** further comprising populating a library of predetermined keywords with names corresponding to playlists designated as favorites by said one or more user profiles in which the media playback system is registered, the method according to any one of claims 1 to 12. **Claim 14** further comprising determining that a first playback command requires a given parameter, playing an audio prompt to provide a second voice input including a keyword corresponding to the given parameter, wherein the second voice input includes a second voice utterance, determining, via the local NLU, whether the second voice utterance includes the keyword corresponding to the given parameter, and executing the first playback command according to the given parameter if the second voice utterance includes the keyword corresponding to the given parameter, the method according to any one of claims 1 to 13. **Claim 15** A non-transitory computer-readable medium storing instructions executable by one or more processors to cause a playback device to execute the method according to any one of claims 1 to 14. **Claim 16** A playback device, comprising a network interface, one or more processors, at least one microphone configured to detect sound, at least one speaker, first and second wake word engines, and data storage storing instructions executable by said one or more processors to cause the playback device to execute the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Reproducing device and program

    JP2004163590A

  • Device, method and program for dialog with user

    JP2008217444A

  • Information security / privacy in an always listening assistant device

    US10002259B1

  • Methods and apparatus for detecting a voice command

    US20140274203A1

  • Contextual hotwords

    US20180182390A1