Linear filtering for speech detection in noise suppression

Through multi-microphone noise suppression technology and Wiener filtering, the problem of poor voice recording quality in noisy environments is solved, and high-precision voice detection and wake-up word recognition are achieved.

CN112424864BActive Publication Date: 2025-09-05SONOS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980047117.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-05-18
Filing Date
2019-05-17
Publication Date
2025-09-05
Estimated Expiration
2039-05-17

AI Technical Summary

Technical Problem

In environments with multiple people talking, noise from household appliances, and other extraneous sounds, it is difficult to obtain high-quality voice command recordings for effective analysis.

Method used

Multi-microphone noise suppression technology is adopted, and Wiener filtering is used for linear time-invariant filtering to estimate the noise content and filter out the noise, retaining the speech content. Beamforming technology is combined for spatial positioning to enhance the accuracy of speech detection.

Benefits of technology

The accuracy of voice input detection, especially the accuracy of wake-up word detection, is improved, noise interference is reduced, and the quality of voice processing is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112424864B_ABST
    Figure CN112424864B_ABST
Patent Text Reader

Abstract

A system and method for suppressing noise and detecting voice input in a multi-channel audio signal captured by multiple microphones includes: (i) capturing a first audio signal via a first microphone and a second audio signal via a second microphone, wherein the first audio signal and the second audio signal respectively include first noise content and second noise content from a noise source; (ii) identifying the first noise content in the first audio signal; (iii) using the identified first noise content to determine estimated noise content captured by the multiple microphones; (iv) using the estimated noise content to suppress the first noise content and the second noise content in the first audio signal and the second audio signal; (v) combining the suppressed first audio signal and the second audio signal into a third audio signal; and (vi) determining that the third audio signal includes voice input including a wake-up word.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Patent Application No. 15 / 984,073, filed May 18, 2018, which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure relates to consumer products and, more particularly, to methods, systems, products, features, services, and other elements directed to media playback and aspects thereof. Background Art

[0004] Options for accessing and listening to digital audio from a loudspeaker setup were limited until 2003, when Sonos filed one of its first patent applications, titled "Method for Synchronizing Audio Playback between Multiple Networked Devices," and began selling media playback systems in 2005. The Sonos wireless high-fidelity (HiFi) system enables people to experience music from many sources through one or more networked playback devices. Using a software control application installed on a smartphone, tablet, or computer, a person can play the content they want in any room with a networked playback device. In addition, using a controller, for example, different songs can be streamed to each room with a playback device, rooms can be grouped together for synchronized playback, or the same song can be listened to simultaneously in all rooms.

[0005] Given the growing interest in digital media, there remains a need to develop technology that is easy for consumers to use to further enhance the listening experience. Summary of the Invention

[0006] This disclosure describes systems and methods for, among other things, processing audio content captured by multiple networked microphones to suppress noise content from the captured audio and detect speech input in the captured audio.

[0007] Some example embodiments relate to capturing, via multiple microphones of a network device, (i) a first audio signal via a first microphone of the multiple microphones, and (ii) a second audio signal via a second microphone of the multiple microphones. The first audio signal includes a first noise content from a noise source, and the second audio signal includes a second noise content from the same noise source. The network device identifies the first noise content in the first audio signal and uses the identified first noise content to determine an estimated noise content captured by the multiple microphones. The network device then uses the estimated noise content to suppress the first noise content in the first audio signal and the second noise content in the second audio signal. The network device combines the suppressed first audio signal and the suppressed second audio signal into a third audio signal. Ultimately, the network device determines that the third audio signal includes a voice input including a wake-up word, and in response to the determination, sends at least a portion of the voice input to a remote computing device for voice processing to identify a voice utterance different from the wake-up word.

[0008] Some embodiments include an article of manufacture comprising a tangible, non-transitory computer-readable medium storing program instructions that, upon execution by one or more processors of a network device, cause the network device to perform operations according to example embodiments disclosed herein.

[0009] Some embodiments include a network device comprising one or more processors and a tangible, non-transitory computer-readable medium storing program instructions that, when executed by the one or more processors, cause the network device to perform operations according to the example embodiments disclosed herein.

[0010] This summary is illustrative only and is not intended to be limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The features, aspects, and advantages of the disclosed technology may be better understood with reference to the following description, appended claims, and accompanying drawings, in which:

[0012] Figure 1 shows an example media playback system configuration in which certain embodiments may be practiced;

[0013] Figure 2 shows a functional block diagram of an example playback device;

[0014] Figure 3 shows a functional block diagram of an example control device;

[0015] Figure 4 An example controller interface is shown;

[0016] Figure 5 An example plurality of network devices is shown;

[0017] Figure 6 A functional block diagram of an example network microphone device is shown;

[0018] Figure 7A An example network device with microphones arranged in a beamforming array is shown in accordance with some embodiments.

[0019] Figure 7B An example network device with microphones arranged in an unordered manner is shown in accordance with some embodiments.

[0020] Figure 7C Shown are two example network devices with a microphone arranged between the two devices in accordance with some embodiments.

[0021] Figure 8 An example media network configuration is shown in which certain embodiments may be practiced.

[0022] Figure 9 Example methods according to some embodiments are shown.

[0023] Figure 10 An example speech input is shown in accordance with some embodiments.

[0024] Figure 11 Experimental results showing improvement in wake-up word detection through static beamforming techniques are shown.

[0025] The drawings are for the purpose of illustrating example embodiments, it being understood that the present invention is not limited to the arrangements and instrumentality shown. DETAILED DESCRIPTION

[0026] I. Overview

[0027] This disclosure describes systems and methods for, among other things, performing noise suppression using a network of microphones. In some embodiments, one or more microphones of a microphone network are components of a network device, such as a voice-enabled device ("VED"). In operation, a microphone-equipped VED (or other network device) listens for a "wake-up word" or wake-up phrase that prompts the VED to capture speech for voice command processing. In some embodiments, the wake-up phrase includes the wake-up word, and vice versa.

[0028] Some examples of a “wake-up word” (or wake-up phrase) may include: “Hey, Sonos” for a Sonos VED, “Alexa” for an Amazon VED, or “Siri” for an Apple VED. Other VEDs from other manufacturers may use different wake-up words and / or wake-up phrases. In operation, a VED equipped with a microphone listens for its wake-up word. And in response to detecting its wake-up word, the VED (alone or in combination with one or more other computing devices) records speech following the wake-up word, analyzes the recorded speech to determine the voice command, and then implements the voice command. Typical voice command examples include: “Play my Beatles playlist,” “Turn on my living room lights,” “Set my thermostat to 75 degrees,” “Add milk and bananas to my shopping list,” and so on.

[0029] Figure 10 An example of voice input 1090 that may be provided to a VED is shown. The voice input 1090 may include a wake word 1092, a voice utterance 1094, or both. The voice utterance portion 1094 may include, for example, one or more spoken commands 1096 (identified as a first command 1096a and a second command 1096b, respectively) and one or more spoken keywords 1098 (identified as a first keyword 1098a and a second keyword 1098b, respectively). In one example, the first command 1096a may be a command to play music, such as a specific song, album, playlist, etc. In this example, the keyword 1098 may be a keyword that identifies one or more regions (e.g., a region) in which the music is to be played. Figure 1 In some examples, the speech utterance portion 1094 may include other information, such as detected pauses (e.g., periods of non-speech) between words spoken by the user, such as Figure 10 Pauses can distinguish within speech utterance 1094 where individual commands, keywords, or other information are spoken by the user.

[0030] like Figure 10 As further shown, the VED can instruct the playback device to temporarily reduce the amplitude of the audio content playback (or "cancel") during the capture of the wake word and / or voice utterance 1096 including the command. Cancel can reduce audio interference and improve voice processing accuracy. Various examples of wake words, voice commands, and related voice input capture techniques, processes, devices, and systems can be found, for example, in U.S. patent application Ser. No. 15 / 721,141, filed on Sep. 27, 2017, entitled "Media Playback System with Voice Assistance," the entire contents of which are incorporated herein by reference.

[0031] One challenge in determining voice commands is obtaining a high-quality recording of speech, including voice commands, for analysis. A higher-quality recording of speech, including voice commands, makes it easier for speech algorithms to analyze than a lower-quality recording of speech, including voice commands. Obtaining a high-quality recording of speech, including voice commands, can be challenging in environments where there may be multiple people speaking, noise from household appliances (e.g., televisions, stereos, air conditioners, dishwashers, etc.), and other extraneous sounds.

[0032] One method of improving the quality of sound recordings including voice commands is to employ a microphone array and use beamforming to (i) amplify the sound from the direction in which the speech containing the voice command is oriented relative to the microphone array, and (ii) attenuate the sound from other directions relative to the microphone array. In a beamforming system, a plurality of microphones arranged in a structured array can perform spatial localization of the sound relative to the microphone array (i.e., determine the direction from which the sound originates). However, although it effectively suppresses unnecessary noise from sound recordings, beamforming has limitations. For example, since beamforming requires the microphones to be arranged in a specific array configuration, beamforming is only feasible in scenarios where such a microphone array can be implemented. Due to hardware or other design constraints, some network devices may not be able to support such a microphone array. As described in more detail below, network devices configured according to various embodiments of the technology and associated systems and methods can address these and other challenges associated with conventional techniques (e.g., traditional beamforming) to suppress noise content from the captured audio.

[0033] This disclosure describes the use of multi-microphone noise suppression techniques that do not necessarily rely on the geometric arrangement of microphones. Instead, techniques for suppressing noise according to various embodiments include: assuming known fixed signal and noise spectra, linear time-invariant filtering of the observed noise process and additive noise. In some embodiments, the present technology uses first audio content captured by one or more corresponding microphones within a microphone network to estimate the noise in second audio content simultaneously captured by one or more other corresponding microphones of the microphone network. The estimated noise from the first audio content can then be used to filter out the noise and preserve the speech in the second audio content.

[0034] In various embodiments, the present technology may involve aspects of Wiener filtering. Conventional Wiener filtering techniques have been used for image filtering and noise removal, but typically include concerns about the fidelity of the resulting filtered signal. However, the inventors have recognized that related techniques based on Wiener filtering can be applied to voice input detection (e.g., wake-up word detection) in a manner that enhances voice detection accuracy compared to voice input detection using conventional beamforming techniques.

[0035] In some embodiments, the microphone network implementing the multi-microphone noise suppression techniques of various embodiments is a component of a network device. A network device is any computing device that includes (i) one or more processors, (ii) one or more network interfaces and / or one or more other types of communication interfaces, and (iii) a tangible, non-transitory computer-readable medium containing instructions encoded therein, wherein the instructions, when at least partially executed by the one or more processors, cause the network device to perform the functions disclosed and described herein. A network device is a general class of devices that includes, but is not limited to, voice-enabled devices (VEDs), networked microphone devices (NMDs), audio playback devices (PBDs), and video playback devices (VPDs). A VED is a class of devices that includes, but is not limited to, NMDs, PBDs, and VPDs. For example, one type of VED is an NMD, which is a network device that includes one or more processors, a network interface, and one or more microphones. Some NMDs may additionally include one or more speakers and perform media playback functions. Another type of VED is a PBD, which is a network device that includes one or more processors, a network interface, and one or more speakers. Some PBDs may optionally include one or more microphones and perform the functions of an NMD. Another type of VED is a VPD, which is a network device that includes one or more processors, a network interface, one or more speakers, and at least one video display. Some VPDs may optionally include one or more microphones and perform the functions of an NMD. PBDs and VPDs can generally be referred to as media playback devices.

[0036] Each of the aforementioned VEDs may implement at least some voice control functionality that allows the VED (alone or possibly in combination with one or more other computing devices) to act on voice commands received via its microphone, thereby allowing a user to control the VED and (possibly) other devices.

[0037] Other embodiments include tangible, non-transitory computer-readable media having stored thereon program instructions that, when executed by a computing device, cause the computing device to perform the features and functions disclosed and described herein.

[0038] Some embodiments include a computing device comprising at least one processor, a data storage device, and program instructions. In operation, the program instructions are stored in the data storage device and, when executed by the at least one processor, cause the computing device (alone or in combination with other components or systems) to perform the features and functions disclosed and described herein.

[0039] Although some examples described herein may involve functions being performed by given actors (e.g., "users" and / or other entities), it should be understood that this is for purposes of explanation only. The claims should not be interpreted as requiring any such example actors to perform an action unless the language of the claims themselves expressly requires it. Those of ordinary skill in the art will appreciate that the disclosure encompasses many other embodiments.

[0040] II. Sample Operating Environment

[0041] Figure 1 1 shows an example configuration of a media playback system 100 in which one or more embodiments disclosed herein may be practiced or implemented. The media playback system 100 as shown is associated with an example home environment having several rooms and spaces (e.g., a master bedroom, a study, a dining room, and a living room). Figure 1 As shown in the example of , media playback system 100 includes playback devices 102-124, control devices 126 and 128, and a wired or wireless network router 130. In operation, any of playback devices (PBD) 102-124 may be a voice-enabled device (VED) as previously described.

[0042] Further discussion of the different components of the example media playback system 100 and how the different components may interact to provide a media experience to a user may be found in the following sections. Although the discussion herein may generally relate to the example media playback system 100, the techniques described herein are not limited to, in particular, Figure 1 For example, the technology described herein may be useful in environments where multi-zone audio may be desired, such as commercial environments such as restaurants, shopping malls, or airports, vehicles such as sport utility vehicles (SUVs), buses or cars, ships or boats, airplanes, and the like.

[0043] a. Example playback device

[0044] Figure 2 A functional block diagram of an example playback device 200 is shown, which may be configured as Figure 1

[0026] One or more of the playback devices 102-124 of the media playback system 100. As mentioned above, the playback device (PBD) 200 is a type of voice-enabled device (VED).

[0045] The playback device 200 includes one or more processors 202, software components 204, memory 206, an audio processing component 208, an audio amplifier 210, a speaker 212, a network interface 214 including a wireless interface 216 and a wired interface 218, and a microphone 220. In one embodiment, the playback device 200 may not include the speaker 212, but may include a speaker interface for connecting the playback device 200 to an external speaker. In another embodiment, the playback device 200 may include neither the speaker 212 nor the audio amplifier 210, but may include an audio interface for connecting the playback device 200 to an external audio amplifier or audio-visual receiver.

[0046] In some examples, one or more processors 202 include one or more clock-driven computing components that are configured to process input data according to instructions stored in memory 206. Memory 206 can be a tangible, non-transitory computer-readable medium that is configured to store instructions that can be executed by one or more processors 202. For example, memory 206 can be a data storage device that can be loaded with one or more software components 204 that can be executed by one or more processors 202 to implement specific functions. In one example, these functions can involve playback device 200 retrieving audio data from an audio source or another playback device. In another example, these functions can involve playback device 200 sending audio data to another device or playback device on a network. In yet another example, these functions can involve playback device 200 pairing with one or more playback devices to create a multi-channel audio environment.

[0047] Particular functionality may involve synchronizing playback of audio content by the playback device 200 with one or more other playback devices. During synchronized playback, a listener will preferably be unable to perceive a time delay difference between the playback of the audio content by the playback device 200 and the one or more other playback devices. Some examples of audio playback synchronization between playback devices are provided in more detail in U.S. Patent No. 8,234,395, entitled “System and method for synchronizing operations among a plurality of independently clocked digital data processing devices,” which is incorporated herein by reference.

[0048] The memory 206 may also be configured to store data associated with the playback device 200, such as one or more regions and / or region groups of which the playback device 200 is a part, audio sources accessible to the playback device 200, or playback queues with which the playback device 200 (or some other playback device) may be associated. The data may be stored as one or more state variables that are periodically updated and used to describe the state of the playback device 200. The memory 206 may also include data associated with the state of other devices of the media system and may be occasionally shared between devices so that one or more of the devices have the latest data associated with the system. Other embodiments are also possible.

[0049] The audio processing component 208 may include one or more digital-to-analog converters (DACs), audio pre-processing components, audio enhancement components, or digital signal processors (DSPs), among others. In one embodiment, one or more of the audio processing components 208 may be subcomponents of one or more processors 202. In one example, the audio processing component 208 may process and / or intentionally alter audio content to generate an audio signal. The generated audio signal may then be provided to an audio amplifier 210 for amplification and playback through a speaker 212. Specifically, the audio amplifier 210 may include a device configured to amplify the audio signal to a level sufficient to drive one or more of the speakers 212. The speaker 212 may include a separate transducer (e.g., a "driver") or a complete speaker system including an enclosure having one or more drivers. For example, a specific driver for a speaker 212 may include, for example, a woofer (e.g., for low frequencies), a mid-range driver (e.g., for mid-range frequencies), and / or a tweeter (e.g., for high frequencies). In some cases, each transducer in one or more speakers 212 may be driven by a corresponding audio amplifier in the audio amplifier 210. In addition to generating analog signals for playback by playback device 200 , audio processing component 208 may also be configured to process audio content to be sent to one or more other playback devices for playback.

[0050] Audio content to be processed and / or played back by the playback device 200 may be received from an external source, for example, via an audio line-in input connection (eg, an auto-detecting 3.5 mm audio line-in connection) or the network interface 214 .

[0051] The network interface 214 can be configured to facilitate the flow of data between the playback device 200 and one or more other devices on the data network, including but not limited to data to / from other VEDs (e.g., commands to perform SPL measurements, SPL measurement data, commands to set system response quantities, and other data and / or commands that facilitate the execution of the features and functions disclosed and described herein). In this way, the playback device 200 can be configured to receive audio content via the data network from one or more other playback devices in communication with the playback device 200, network devices within a local area network, or audio content sources on a wide area network (e.g., the Internet). The playback device 200 can send metadata to and / or receive metadata from other devices on the network, including but not limited to components of the networked microphone system disclosed and described herein. In one example, the audio content and other signals (e.g., metadata and other signals) sent and received by the playback device 200 can be sent in the form of digital packet data that includes an Internet Protocol (IP)-based source address and an IP-based destination address. In this case, the network interface 214 may be configured to parse the digital packet data so that the data destined for the playback device 200 is properly received and processed by the playback device 200 .

[0052] As shown, the network interface 214 may include a wireless interface 216 and a wired interface 218. The wireless interface 216 may provide a network interface function for the playback device 200 to wirelessly communicate with other devices (e.g., other playback devices, speakers, receivers, network devices, control devices within a data network associated with the playback device 200) according to a communication protocol (e.g., any wireless standard including IEEE802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G mobile communication standards, etc.). The wired interface 218 may provide a network interface function for the playback device 200 to communicate with other devices via a wired connection according to a communication protocol (e.g., IEEE 802.3). Although Figure 2 The network interface 214 shown in FIG. 2 includes both a wireless interface 216 and a wired interface 218 , but in some embodiments, the network interface 214 may include only a wireless interface or only a wired interface.

[0053] Microphone 220 can be arranged to detect sounds in the environment of playback device 200. For example, the microphone can be mounted on the outer wall of the housing of the playback device. The microphone can be any type of microphone now known or later developed, for example, a condenser microphone, an electret condenser microphone, or a dynamic microphone. The microphone can be sensitive to a portion of the frequency band of speaker 220. One or more of speakers 220 can operate inversely to microphone 220. In some aspects, playback device 200 may not have microphone 220.

[0054] In one example, playback device 200 and another playback device can be paired to play two separate audio components of audio content. For example, playback device 200 can be configured to play the left channel audio component, while another playback device can be configured to play the right channel audio component, thereby producing or enhancing the stereo effect of the audio content. Paired playback devices (also referred to as "bound playback devices") can also play audio content synchronously with other playback devices.

[0055] In another example, playback device 200 can merge with one or more other playback device sounds to form a single merged playback device. The merged playback device can be configured to process and reproduce sound differently from the playback device of non-merged or paired playback device, because the merged playback device can have an additional speaker driver that can present audio content by it. For example, if playback device 200 is a playback device (that is, a subwoofer) designed to present low-frequency band audio content, playback device 200 can merge with the playback device designed to present full-band audio content. In this case, when merging with low-frequency playback device 200, full-band playback device can be configured to only present the mid-high frequency component of audio content, while low-frequency band playback device 200 then presents the low-frequency component of audio content. The merged playback device can also be paired with a single playback device or another merged playback device.

[0056] For example, SONOS currently offers (or has offered) for sale certain playback devices including "PLAY:1," "PLAY:3," "PLAY:5," "PLAYBAR," "CONNECT:AMP," "CONNECT," and "SUB."

[0057] Any other past, present and / or future playback devices may additionally or alternatively be used to implement the playback devices of the example embodiments disclosed herein. Furthermore, it should be understood that the playback devices are not limited to Figure 2Examples are shown or Sonos' product offerings. For example, the playback device can include wired or wireless headphones. In another example, the playback device can include or interact with a docking station for a personal mobile media playback device. In yet another example, the playback device can be integrated into another device or component, such as a television, lighting fixture, or some other device used indoors or outdoors.

[0058] b. Example playback region configuration

[0059] Return Reference Figure 1 The media playback system 100 may have one or more playback zones, each of which may have one or more playback devices and / or other VEDs. The media playback system 100 may be established with one or more playback zones, and one or more zones may be added or removed later to achieve the desired playback performance. Figure 1 Example configuration shown. Each zone can be named after a different room or space (e.g., study, bathroom, master bedroom, bedroom, kitchen, dining room, living room, and / or balcony). In one case, a single playback zone can include multiple rooms or spaces. In another case, a single room or space can include multiple playback zones.

[0060] like Figure 1 As shown, the balcony, dining room, kitchen, bathroom, study, and bedroom areas each have one playback device, while the living room and master bedroom areas each have multiple playback devices. In the living room area, playback devices 104, 106, 108, and 110 can be configured to play audio content synchronously as individual playback devices, as one or more bundled playback devices, as one or more combined playback devices, or any combination thereof. Similarly, in the master bedroom, playback devices 122 and 124 can be configured to play audio content synchronously as individual playback devices, as one or more bundled playback devices, or as one or more combined playback devices.

[0061] In one example, Figure 1One or more playback areas in an environment can all play different audio content. For example, a user can be grilling in the balcony area and listening to hip-hop music being played by playback device 102, while another user can be preparing food in the kitchen area and listening to classical music being played by playback device 114. In another example, a playback area can play the same audio content synchronously with another playback area. For example, a user can be in the study area, where playback device 118 is playing the same rock music as the rock music being played by playback device 102 in the balcony area. In this case, playback devices 102 and 118 can play rock music synchronously, so that the user can seamlessly (or at least substantially seamlessly) enjoy the audio content being played when moving between different playback areas. Synchronization between playback areas can be achieved in a manner similar to the synchronization between playback devices described in previously cited U.S. Patent No. 8,234,395.

[0062] As suggested above, the zone configuration of the media playback system 100 can be modified dynamically, and in some embodiments, the media playback system 100 supports multiple configurations. For example, if a user physically moves one or more playback devices into or out of a zone, the media playback system 100 can be reconfigured to accommodate the change. For example, if a user physically moves playback device 102 from the balcony zone to the study zone, the study zone can now include both playback device 118 and playback device 102. Playback device 102 can be paired or grouped with the study zone and / or renamed (if desired) via control devices (e.g., control devices 126 and 128). On the other hand, if one or more playback devices are moved to a specific area of ​​the home environment that is not already a playback zone, a new playback zone can be created for that specific area.

[0063] Furthermore, different playback zones of the media playback system 100 can be dynamically combined into zone groups or divided into separate playback zones. For example, the dining room zone and the kitchen zone can be combined into a zone group for a banquet, so that playback devices 112 and 114 can present (e.g., play back) audio content synchronously. On the other hand, if a user wants to listen to music in the living room space while another user wants to watch TV, the living room zone can be divided into a TV zone including playback device 104 and a listening zone including playback devices 106, 108, and 110.

[0064] c. Example Control Device

[0065] Figure 31 shows a functional block diagram of an example control device 300 that can be configured as one or both of the control devices 126 and 128 of the media playback system 100. As shown, the control device 300 can include one or more processors 302, a memory 304, a network interface 306, a user interface 308, a microphone 310, and a software component 312. In one example, the control device 300 can be a dedicated controller for the media playback system 100. In another example, the control device 300 can be a network device, such as an iPhone, on which the media playback system controller application software can be installed. TM , iPad TM or any other smartphone, tablet or web-connected device (e.g., an internet-connected computer such as a PC or Mac TM ).

[0066] The one or more processors 302 can be configured to perform functions related to facilitating user access, control, and configuration of the media playback system 100. The memory 304 can be a data storage device that can be loaded with one or more software components executable by the one or more processors 302 for performing these functions. The memory 304 can also be configured to store media playback system controller application software and other data associated with the media playback system 100 and the user. In one example, the network interface 306 can be based on industry standards (e.g., infrared; radio; wired standards including IEEE 802.3; wireless standards including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 3G, 4G, or 5G mobile communication standards, etc.). The network interface 306 can provide a means for the control device 300 to communicate with other devices in the media playback system 100. In one example, data and information (e.g., state variables) can be transferred between the control device 300 and the other devices via the network interface 306. For example, playback zone and zone group configurations in the media playback system 100 may be received by the control device 300 from a playback device or another network device via the network interface 306, or sent by the control device 300 to another playback device or network device via the network interface 306. In some cases, the other network device may be another control device.

[0067] Playback device control commands (e.g., volume control and audio playback control) may also be transmitted from the control device 300 to the playback device via the network interface 306. As suggested above, the control device 300 may also be used by a user to perform changes to the configuration of the media playback system 100. Configuration changes may include: adding / removing one or more playback devices to / from a zone; adding / removing one or more zones to / from a zone group; forming a bound or merged player; detaching one or more playback devices from a bound or merged player; and the like. Thus, the control device 300 may sometimes be referred to as a controller, regardless of whether the control device 300 is a dedicated controller or a network device having the media playback system controller application software installed thereon.

[0068] The control device 300 may include a microphone 310. The microphone 310 may be arranged to detect sounds in the environment of the control device 300. The microphone 310 may be any type of microphone now known or later developed, such as a condenser microphone, an electret condenser microphone, or a dynamic microphone. The microphone may be sensitive to a portion of the frequency band. Two or more microphones 310 may be arranged to capture location information of an audio source (e.g., speech, audible sound) and / or to help filter background noise.

[0069] The user interface 308 of the control device 300 may be configured to provide a controller interface (e.g., Figure 4 The example controller interface 400 shown in FIG. 4 is used to facilitate user access and control of the media playback system 100. The controller interface 400 includes a playback control area 410, a playback zone area 420, a playback status area 430, a playback queue area 440, and an audio content source area 450. The user interface 400 shown is only an example of a user interface that can be used on a network device (e.g., Figure 3 The control device 300 (and / or Figure 1 The present invention is one example of a user interface provided on control devices 126 and 128 of the network and accessed by a user to control a media playback system (e.g., media playback system 100). Alternatively, other user interfaces of varying formats, styles, and interaction sequences may be implemented on one or more network devices to provide similar control access to the media playback system.

[0070] The playback control area 410 may include selectable (e.g., by touch or by using a cursor) icons to cause the playback device in the selected playback zone or zone group to play or pause, fast forward, rewind, skip to the next track, skip to the previous track, enter / exit shuffle mode, enter / exit repeat mode, enter / exit crossfade mode, etc. The playback control area 410 may also include selectable icons for modifying equalization settings, playback volume, etc.

[0071] The playback zone area 420 may include representations of playback zones within the media playback system 100. In some embodiments, the graphical representations of playback zones may be selectable to bring up additional selectable icons for managing or configuring playback zones in the media playback system, e.g., creating bound zones, creating zone groups, detaching zone groups, renaming zone groups, etc.

[0072] For example, as shown in the figure, a "grouping" icon can be provided in each graphical representation of a playback area. The "grouping" icon provided in the graphical representation of a specific area can be selectable, so as to call out the option for selecting one or more other areas in the media playback system that will be grouped with the specific area. Once grouped, the playback devices in the area that has been grouped with the specific area will be configured to play audio content synchronously with the playback devices in the specific area. Similarly, a "grouping" icon can be provided in the graphical representation of a regional group. In this case, the "grouping" icon can be selectable, to call out the option for canceling the selection of one or more areas to be removed from the regional group. Other interactions and implementations of grouping and canceling groups of areas via a user interface (e.g., user interface 400) are also possible. When the playback area or regional group configuration is modified, the representation of the playback area in the playback area area 420 can be dynamically updated.

[0073] The playback status area 430 may include a graphical representation of the audio content currently playing, previously playing, or scheduled to play next in the selected playback zone or zone group. The selected playback zone or zone group may be visually distinguished on the user interface, for example, within the playback zone area 420 and / or the playback status area 430. The graphical representation may include the track name, artist name, album name, album year, track length, and other relevant information that the user may find useful when controlling the media playback system via the user interface 400.

[0074] Playback queue area 440 can comprise the graphical representation of the audio content in the playback queue that is associated with selected playback area or regional group.In certain embodiments, each playback area or regional group can be associated with playback queue, and this playback queue comprises the corresponding information with zero or more audio items played back by this playback area or regional group.For example, each audio item in the playback queue can comprise uniform resource identifier (URI), uniform resource locator (URL) or some other identifiers, it can be used for searching and / or retrieving audio items from local audio content source or networked audio content source by the playback device in the playback area or regional group, may play back for playback device.

[0075] In one example, playlist can be added to playback queue, in which case the information corresponding to each audio item in the playlist can be added to playback queue. In another example, the audio items in the playback queue can be saved as playlist. In another example, when playback area or regional group are continuously playing streaming audio content (for example, internet radio, which can be continuously played until stopped), rather than with the discrete audio items of playback duration, playback queue can be empty or filled but "unused". In an alternative embodiment, playback queue can comprise internet radio and / or other streaming audio content items, and when playback area or regional group are playing these items, are in "use". Other examples are also possible.

[0076] When playback zones or zone groups are "grouped" or "ungrouped," the playback queues associated with the affected playback zones or zone groups can be cleared or reassociated. For example, if a first playback zone including a first playback queue is grouped with a second playback zone including a second playback queue, the established zone group can have an associated playback queue that is initially empty, contains audio items from the first playback queue (e.g., if the second playback zone was added to the first playback zone), or contains audio items from the second playback queue (e.g., if the first playback zone was added to the second playback zone), or contains a combination of audio items from both the first playback queue and the second playback queue. Subsequently, if the established zone group is ungrouped, the resulting first playback zone can be reassociated with the previous first playback queue, or associated with a new playback queue that is empty or contains audio items from a playback queue associated with the zone group that was established before the established zone group was ungrouped. Similarly, the resulting second playback zone can be reassociated with the previous second playback queue, or associated with a new playback queue that is empty or contains audio items from a playback queue associated with a zone group that was established before the zone group was ungrouped. Other examples are possible.

[0077] Return Reference Figure 4In the user interface 400 of FIG. 4 , the graphical representation of audio content in the playback queue area 440 may include track title, artist name, track length, and other relevant information associated with the audio content in the playback queue. In one example, the graphical representation of audio content may be selectable to bring out additional selectable icons to manage and / or manipulate the playback queue and / or the audio content represented in the playback queue. For example, the represented audio content may be removed from the playback queue, moved to a different position in the playback queue, or selected to play immediately, or after any currently playing audio content, etc. The playback queue associated with a playback region or region group may be stored in a memory on one or more playback devices in the playback region or region group, on a playback device not in the playback region or region group, and / or on some other designated device.

[0078] The audio content source area 450 may include graphical representations of selectable audio content sources from which audio content may be obtained and played by the selected playback zone or zone group. A discussion of audio content sources may be found in the following section.

[0079] d. Sample audio content source

[0080] As previously mentioned, one or more playback devices in a region or region group can be configured to retrieve playback audio content (e.g., according to the corresponding URI or URL of the audio content) from various available audio content sources. In one example, the playback device can directly retrieve the audio content from a corresponding audio content source (e.g., a line-in connection). In another example, audio content can be provided to the playback device over a network, via one or more other playback devices or network devices.

[0081] Example audio content sources may include: a media playback system (e.g., Figure 1 The audio content may be from a memory of one or more playback devices in the media playback system 100, a local music library on one or more network devices (e.g., a controller device, a network-enabled personal computer, or a network attached storage (NAS)), a streaming audio service that provides audio content over the Internet (e.g., the cloud), or an audio source connected to the media playback system via a line-in connection on a playback device or network device.

[0082] In some embodiments, the media playback system (e.g., Figure 1In one embodiment, the audio content source is regularly added to the media playback system 100, or the audio content source is removed therefrom. In one example, whenever adding, removing or updating one or more audio content sources, audio item indexing can be performed. Audio item indexing can be comprised by: scanning the identifiable audio items in all folders / directories shared on the accessible network of the playback device in the media playback system, and generating or updating the audio content database that comprises metadata (for example, title, artist, album, track length etc.) and other associated information (for example, URL or the URI for each identifiable audio item found). Other examples for managing and maintaining the audio content source are also possible.

[0083] The above discussion of playback devices, controller devices, playback zone configurations, and media content sources provides only some examples of operating environments in which the functions and methods described below may be implemented. Configurations of media playback systems, playback devices, and network devices not explicitly described herein and other operating environments may also be applicable and suitable for implementation of the functions and methods.

[0084] e. Example of multiple network devices

[0085] Figure 5 An example plurality of network devices 500 are shown that can be configured to provide an audio playback experience using voice control. One of ordinary skill in the art will appreciate that Figure 5 The devices shown in FIG are for illustrative purposes only, and variations including different and / or additional (or fewer) devices are possible. As shown, multiple network devices 500 include computing devices 504, 506, and 508; network microphone devices (NMDs) 512, 514, 516, and 518; playback devices (PBDs) 532, 534, 536, and 538; and controller device 522. As previously described, any one or more (or all) of NMDs 512-16, PBDs 532-38, and / or controller device 522 may be VEDs. For example, in some embodiments, PBDs 532 and 536 may be VEDs, while PBDs 534 and 538 may not be VEDs.

[0086] Each of the plurality of network devices 500 is a network-capable device that can communicate with a user according to one or more network protocols (e.g., NFC, Bluetooth, TM , Ethernet, and IEEE 802.11, etc.), establish communication with one or more other devices in multiple devices on one or more types of networks (e.g., wide area network (WAN), local area network (LAN), and personal area network (PAN), etc.).

[0087] As shown, computing devices 504, 506, and 508 are part of cloud network 502. Cloud network 502 may include additional computing devices (not shown). In one example, computing devices 504, 506, and 508 may be different servers. In another example, two or more of computing devices 504, 506, and 508 may be modules of a single server. Similarly, each of computing devices 504, 506, and 508 may include one or more modules or servers. For ease of description herein, each of computing devices 504, 506, and 508 may be configured to perform a specific function within cloud network 502. For example, computing device 508 may be a source of audio content for a streaming music service, while computing device 506 may be connected to a voice assistant service (e.g., a voice assistant service) for processing voice input that has been captured after detecting a wake word. Google or other voice services). As an example, the VED may send the captured voice input (e.g., the voice utterance and the wake word) or a portion thereof (e.g., the voice utterance immediately following the wake word) to the computing device 506 over a data network for speech processing. The computing device 506 may employ a text-to-speech engine to convert the voice input into text, which may be processed to determine the underlying intent of the voice utterance. The computing device 506 or another computing device may send a corresponding response to the voice input to the VED, for example, a response including as its payload one or more audible outputs (e.g., a voice response to a query and / or confirmation) and / or instructions intended for one or more network devices of the local system. The instructions may include, for example, commands for initiating, pausing, resuming, or stopping playback of audio content on one or more network devices, increasing / decreasing playback volume, retrieving tracks or playlists corresponding to an audio queue via a specific URI or URL, and the like. Additional examples of speech processing that determine intent and respond to voice input may be found, for example, in previously referenced U.S. patent application No. 15 / 721,141.

[0088] As shown, computing device 504 may be configured to interface with NMDs 512, 514, and 516 via communication path 542. NMDs 512, 514, and 516 may be components of one or more "smart home" systems. In one embodiment, NMDs 512, 514, and 516 may be physically distributed throughout a home, similar to Figure 1 In another embodiment, two or more of NMDs 512, 514, and 516 may be physically located relatively close to each other. Communication path 542 may include one or more types of networks, such as a WAN, a LAN, and / or a PAN, including the Internet.

[0089] In one example, one or more of NMDs 512, 514, and 516 are devices configured primarily for audio detection. In another example, one or more of NMDs 512, 514, and 516 may be components of devices having various primary utilities. For example, as described above in conjunction with Figure 2 and Figure 3 As discussed, one or more of NMDs 512, 514, and 516 may be (or at least may include) microphone 220 of playback device 200 or microphone 310 of network device 300 (or at least may be a component thereof). Furthermore, in some cases, one or more of NMDs 512, 514, and 516 may be (or at least may include) playback device 200 or network device 300 (or at least may be a component thereof). In an example, one or more of NMDs 512, 514, and / or 516 may include a plurality of microphones arranged in a microphone array. In some embodiments, one or more of NMDs 512, 514, and / or 516 may be a microphone on a mobile computing device (e.g., a smartphone, tablet computer, or other computing device).

[0090] As shown, computing device 506 is configured to interface with controller device 522 and PBDs 532, 534, 536, and 538 via communication path 544. In one example, controller device 522 may be a network device, e.g., Figure 2 Therefore, the controller device 522 can be configured to provide Figure 4 Controller interface 400. Similarly, PBDs 532, 534, 536, and 538 may be playback devices, such as Figure 3 Thus, PBDs 532, 534, 536, and 538 may be physically distributed throughout the home, such as Figure 1 For illustrative purposes, PBDs 536 and 538 are shown as members of bound region 530, while PBDs 532 and 534 are members of their respective regions. As described above, PBDs 532, 534, 536, and 538 can be dynamically bound, grouped, unbound, and ungrouped. Communication path 544 can include one or more types of networks, such as a WAN including the Internet, a LAN, and / or a PAN.

[0091] In one example, like NMDs 512, 514, and 516, controller device 522 and PBDs 532, 534, 536, and 538 may also be components of one or more "smart home" systems. In one instance, PBDs 532, 534, 536, and 538 are located in the same home as NMDs 512, 514, and 516. Furthermore, as suggested above, one or more of PBDs 532, 534, 536, and 538 may be one or more of NMDs 512, 514, and 516. For example, any one or more (or possibly all) of NMDs 512-16, PBDs 532-38, and / or controller device 522 may be voice-enabled devices (VEDs).

[0092] NMDs 512, 514, and 516 may be part of a local area network, and communication path 542 may include an access point linking the local area network of NMDs 512, 514, and 516 to computing device 504 via a WAN (communication path not shown). Likewise, each of NMDs 512, 514, and 516 may communicate with each other via the access point.

[0093] Similarly, the controller device 522 and PBDs 532, 534, 536, and 538 may be part of a local area network and / or a local playback network (as discussed in the previous section), and the communication path 544 may include an access point that links the local area network and / or local playback network of the controller device 522 and PBDs 532, 534, 536, and 538 to the computing device 506 via the WAN. In this way, the controller device 522 and each of the PBDs 532, 534, 536, and 538 may also communicate with each other via the access point.

[0094] In one example, communication paths 542 and 544 can include the same access point. In an example, each of NMDs 512, 514, and 516, controller device 522, and PBDs 532, 534, 536, and 538 can access cloud network 502 via the same access point of the home.

[0095] like Figure 5 As shown, each of the NMDs 512, 514, and 516, the controller device 522, and the PBDs 532, 534, 536, and 538 may also communicate directly with one or more other devices via a communication means 546. The communication means 546, as described herein, may involve and / or include one or more forms of communication between devices over one or more types of networks according to one or more network protocols, and / or may involve communication via one or more other network devices. For example, the communication means 546 may include Bluetooth TM(IEEE 802.15), NFC, Wi-Fi Direct, and / or proprietary wireless, among others.

[0096] In one example, the controller device 522 can be connected to the TM The NMD 514 may communicate with the controller device 522 via another local area network and may communicate with the PBD 534 via Bluetooth. TM Communicates with PBD 536. In yet another example, each of PBDs 532, 534, 536, and 538 can communicate with each other over a local playback network according to a spanning tree protocol while also communicating with controller device 522 over a local area network different from the local playback network. Other examples are also possible.

[0097] In some cases, the communication method between NMDs 512, 514, and 516, controller device 522, and PBDs 532, 534, 536, and 538 may differ (or may change) depending on the type of communication requirements between the devices, network conditions, and / or latency requirements. For example, when NMD 516 is first introduced into a home with PBDs 532, 534, 536, and 538, communication method 546 may be used. In one embodiment, NMD 516 may send identification information corresponding to NMD 516 to PBD 538 via NFC, and in response, PBD 538 may send local area network information to NMD 516 via NFC (or some other form of communication). However, once NMD 516 is deployed in a home, the communication method between NMD 516 and PBD 538 may change. For example, NMD 516 may subsequently communicate with PBD 538 via communication path 542, cloud network 502, and communication path 544. In another example, the NMD and PBD may never communicate via the local communication means 546. In another example, the NMD and PBD may primarily communicate via the local communication means 546. Other examples are possible.

[0098] In the illustrative example, NMDs 512, 514, and 516 can be configured to receive voice input for controlling PBDs 532, 534, 536, and 538. Available control commands may include any of the media playback system controls previously discussed, such as playback volume control, playback transport control, music source selection and grouping, and the like. In one instance, NMD 512 can receive voice input for controlling one or more of PBDs 532, 534, 536, and 538. In response to receiving the voice input, NMD 512 can send the voice input to computing device 504 via communication path 542 for processing. In one example, computing device 504 can convert the voice input into an equivalent text command and parse the text command to identify the command. Computing device 504 can then send the text command to computing device 506, and computing device 506 can then control one or more of PBDs 532-538 to execute the command. In another example, computing device 504 may convert speech input into equivalent text commands and then send the text commands to computing device 506. Computing device 506 may then parse the text commands to identify one or more playback commands, and computing device 506 may then additionally control one or more of PBDs 532-538 to execute the commands.

[0099] For example, if the text command is "Play track 1 by artist 1 from streaming service 1 in region 1," computing device 506 may identify (i) a URL for track 1 by artist 1 available from streaming service 1, and (ii) at least one playback device in region 1. In this example, the URL for track 1 by artist 1 from streaming service 1 may be a URL pointing to computing device 508, and region 1 may be binding region 530. Thus, upon identifying the URL and one or both of PBDs 536 and 538, computing device 506 may send the identified URL to one or both of PBDs 536 and 538 via communication path 544 for playback. In response, one or both of PBDs 536 and 538 may retrieve audio content from computing device 508 based on the received URL and begin playing track 1 by artist 1 from streaming service 1.

[0100] Those skilled in the art will appreciate that the above are merely illustrative examples, and other implementations are possible. In one embodiment, as described above, the operations performed by one or more of the plurality of network devices 500 may be performed by one or more other devices in the plurality of network devices 500. For example, the conversion of voice input to text commands may be performed alternatively, partially, or entirely by another device or devices, such as the controller device 522, the NMD 512, the computing device 506, the PBD 536, and / or the PBD 538. Similarly, the identification of the URL may be performed alternatively, partially, or entirely by another device or devices, such as the NMD 512, the computing device 504, the PBD 536, and / or the PBD 538.

[0101] f. Sample web microphone device

[0102] Figure 6 A functional block diagram of an example network microphone device 600 is shown. The example network microphone device 600 may be configured as Figure 5 6. The network microphone device 600 may include one or more of the NMDs 512, 514, and 516 and / or any of the VEDs disclosed and described herein. As shown, the network microphone device 600 includes one or more processors 602, tangible, non-transitory computer-readable memory 604, a microphone array 606 (e.g., one or more microphones), a network interface 608, a user interface 610, a software component 612, and a speaker 614. One of ordinary skill in the art will appreciate that other network microphone device configurations and arrangements are possible. For example, the network microphone device may alternatively not include the speaker 614, or may have a single microphone instead of the microphone array 606.

[0103] The one or more processors 602 may include one or more processors and / or controllers, which may take the form of general-purpose or special-purpose processors or controllers. For example, the one or more processing units 602 may include a microprocessor, a microcontroller, an application-specific integrated circuit, a digital signal processor, etc. The tangible non-transitory computer-readable memory 604 may be a data storage device that may be loaded with one or more of the software components executable by the one or more processors 602 for performing these functions. Therefore, the memory 604 may include one or more non-transitory computer-readable storage media, examples of which may include: volatile storage media (e.g., random access memory, registers, cache, etc.), and non-volatile storage media (e.g., read-only memory, hard disk drive, solid-state drive, flash memory and / or optical storage device, etc.).

[0104] The microphone array 606 can be a plurality of microphones arranged to detect sounds in the environment of the network microphone device 600. The microphone array 606 can include any type of microphone now known or developed later, for example, a condenser microphone, an electret condenser microphone, or a dynamic microphone, etc. In one example, the microphone array can be arranged to detect audio from one or more directions relative to the network microphone device. The microphone array 606 can be sensitive to a portion of a frequency band. In one example, a first subset of the microphone array 606 can be sensitive to a first frequency band, while a second subset of the microphone array can be sensitive to a second frequency band. The microphone array 606 can also be arranged to capture location information of an audio source (e.g., speech, audible sound) and / or help filter background noise. It is worth noting that in some embodiments, the microphone array can consist of only a single microphone, rather than a plurality of microphones.

[0105] The network interface 608 may be configured to facilitate various network devices (e.g., Figure 5 , controller device 522, PBDs 532-538, computing devices 504-508 in cloud network 502, and other network microphone devices, etc. Thus, network interface 608 can take any suitable form to perform these functions, examples of which may include: an Ethernet interface, a serial bus interface (e.g., FireWire, USB 2.0, etc.), a chipset and antenna suitable for facilitating wireless communication, and / or any other interface that provides wired and / or wireless communication. In one example, network interface 608 can be based on industry standards (e.g., infrared; radio; wired standards including IEEE 802.3; wireless standards including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G mobile communication standards, etc.).

[0106] The user interface 610 of the network microphone device 600 can be configured to facilitate user interaction with the network microphone device. In one example, the user interface 610 can include one or more of physical buttons, a graphical interface provided on a touch-sensitive screen and / or surface, and the like, for the user to provide input directly to the network microphone device 600. The user interface 610 can also include one or more of a light and a speaker 614 to provide visual and / or audio feedback to the user. In one example, the network microphone device 600 can also be configured to play back audio content through the speaker 614.

[0107] III. Example Noise Suppression Systems and Methods

[0108] Figures 7A-7CNetwork devices 700 are depicted (individually identified as network devices 700a-700d). Each network device 700 includes a housing 704 that at least partially encloses certain components of the network device (not shown), such as amplifiers, transducers, processors, and antennas. Network devices 700 also include microphones 702 (individually identified as microphones 702a-g) arranged at various locations on the housing 704. For example, network device 700a includes a structured array of microphones 702. In some embodiments, microphones 702 can be located within and / or exposed through holes in the housing 704. Network device 700a can be configured to Figure 5 One or more of the NMDs 512, 514, and 516 and / or any of the VEDs disclosed and described herein.

[0109] As described above, the embodiments described herein facilitate suppressing noise from audio content captured by multiple microphones to aid in detecting the presence of a wake-up word in the captured audio content. Some noise suppression processes involve single-microphone techniques for suppressing certain frequencies where noise dominates over speech content. However, these techniques can result in significant distortion of the speech content. Other noise suppression processes involve beamforming techniques, where a structured array of microphones is used to capture audio content from specific directions where speech dominates over noise content, rather than from directions where noise dominates over speech content.

[0110] While effective at suppressing unwanted noise when capturing audio content, beamforming has limitations. For example, compared to the enhanced suppression techniques described below, traditional beamforming may often be suboptimal when detecting speech input. For example, Figure 11 It is shown that using the multi-channel Wiener filter (MCWF) algorithm described below significantly improves wake-word detection relative to traditional static beamforming under the same conditions, which involves (1) detecting the wake-word from a noisy sound sample with an SNR of -15dB, (2) playing back the same sample track ("Relax" by Frankie Goes to Hollywood), and (3) using the same NMD for each test (the NMD used for the test has an array of six microphones that are spaced at an appropriate distance from each other for traditional beamforming). Figure 11Results for three different test cases under test conditions are depicted: (1) Graph 1110 depicts sound samples detected without using beamforming or MCWF algorithms, (2) Graph 1120 depicts sound samples detected using beamforming, and (3) Graph 1130 depicts sound samples detected using an MCWF-based algorithm. In each of Graphs 1110, 1120, and 1130, the x-axis represents time, the y-axis corresponds to frequency, and the darkness of the graph represents the intensity of the detected sound sample in dB (wherein the intensity increases with increasing darkness). In addition, in each of Graphs 1110, 1120, and 1130, the wake-up word (in Figure 11 ) starts approximately halfway along the x-axis and ends three-quarters along the x-axis. Comparing graph 1120 with graph 1110, it can be seen that beamforming removes some noise, but a significant amount of noise still remains. However, comparing graph 1130 with graph 1120, it can be seen that the MCWF algorithm removes significantly more noise than beamforming, and thus the wake word can be more easily recognized from the MCWF-filtered sound sample than from the beamformed sound sample.

[0111] Additionally, beamforming typically requires a known array configuration, and the network device 700 selectively captures audio from specific directions relative to the array. Beamforming may only be feasible if such a microphone array 702 can be implemented. For example, if Figure 7A If the microphones 702 and processing components of the network device 700a are configured for conventional beamforming, then at frequencies up to 4 kHz, using conventional alias-free beamforming, the spacing or distance d1 between adjacent microphones 702 would be limited to a theoretical maximum of approximately 4.25 cm. However, due to hardware or other design constraints, some network devices may not be able to support such a closely spaced array of microphones 702. Therefore, when using the enhanced noise suppression techniques described herein, the distance d1 between the microphones 702 in various embodiments may not be reconstructed to such a theoretical maximum.

[0112] For example, Figure 7B A network device 700b is depicted in which microphones 702 are arranged in an unordered manner. As used herein, the term "unordered manner" refers to any arrangement of microphones that is not used as a beamforming array. As such, microphones arranged in an unordered manner may be arranged in any order relative to one another; more conveniently positioned along a housing, such as between speakers, electronics, buttons, and / or other components; and / or arranged in some order, but in an order that does not support (or at least may not support) beamforming. For example, Figure 7BAs shown, microphones 702 appear to be arranged according to a particular geometric configuration, wherein microphones 702a, 702b, 702f, and 702g are arranged in a first horizontal plane, and microphones 702c, 702d, and 702e are arranged in a second horizontal plane. Figure 7B The arrangement of microphones 702 in the embodiment includes some aspect of order, which is also referred to as "disordered" because the microphones 702 are too dispersed from one another to perform beamforming, or at least too dispersed from one another to perform beamforming effectively for the types of voice applications disclosed and described herein. In some embodiments, the minimum distance between two given microphones is greater than 5 cm. For example, the spacing or distance d2 between microphones 702 c and 702 d, or any other set of two or more microphones, can be between 5 cm and 60 cm.

[0113] Figure 7C Depicted are microphones 702 distributed across multiple network devices according to an example embodiment. Specifically, microphones 702c, 702d, and 702e are disposed in housing 704 of network device 700c, and microphones 702a, 702b, 702f, and 702g are disposed in housing 704 of network device 700d. In some embodiments, network devices 700c and 700d are located in the same room (e.g., as separate devices in a home theater configuration), but in different areas of the room. In such embodiments, the spacing or distance between microphones 702 on network devices 700c and 700d (e.g., distance d3 between microphones 702d and 702f) can exceed 60 cm. For example, distance d3 between microphones 702d and 702f, or any other set of two or more microphones disposed on separate network devices, can be between 1 meter and 5 meters.

[0114] exist Figures 7A-7CIn each of the arrangements depicted in FIG, network device 700 employs a multi-microphone noise suppression technique that does not necessarily rely on the geometric arrangement of microphones 702. Alternatively, techniques for suppressing noise according to various embodiments include linear time-invariant filtering of the observed noise process and additive noise, assuming known fixed signal and noise spectra. Network device 700 uses first audio content captured by one or more microphones 702 to estimate noise in second audio content simultaneously captured by one or more other microphones in microphones 702. For example, microphone 702a captures first audio content, while microphone 702g simultaneously captures second audio content. If a user near network device 700 speaks a voice command, the speech content in both the first audio content captured by microphone 702a and the second audio content captured by microphone 702g includes the same voice command. Furthermore, if a noise source is near network device 700, both the first audio content captured by microphone 702a and the second audio content captured by microphone 702g include noise content from the noise source.

[0115] However, because microphones 702a and 702g are spaced apart from each other, the intensity of the speech content and the noise content can vary between the first audio content and the second audio content. For example, if microphone 702a is closer to the noise source and microphone 702g is closer to the user speaking, the noise content may dominate the first audio content captured by microphone 702a, while the speech content may dominate the second audio content captured by microphone 702g. Furthermore, if the noise content dominates the first audio content, network device 700 can use the first audio content to generate an estimate of the noise content present in the second audio content. The estimated noise from the first audio content can then be used to filter out the noise and retain the speech in the second audio content.

[0116] In some embodiments, network device 700 performs this process simultaneously for all microphones 702, so that the noise content captured by each microphone is used to estimate the noise content captured by each other microphone. Network device 700 uses the estimated noise content to filter the individual audio signals captured by each microphone 702 to suppress the individual noise content in each audio signal, and then combines the filtered audio signals. When the noise content of each audio signal is suppressed, the primary content of each audio signal is voice content, so the combined audio signal is also speech-dominated.

[0117] The following combination Figure 8 An example MCWF algorithm for performing these processes is described in further detail.

[0118] Figure 8An example environment 800 is depicted in which such a noise suppression process is performed. The environment 800 includes a plurality of microphones 802 (identified individually as microphones 802a-g) for capturing audio content. The microphones 802 may be configured to Figures 7A-7C 8. As shown, environment 800 includes seven microphones 802, but in other embodiments, environment 800 includes additional or fewer microphones. In some embodiments, microphones 802 are disposed on or within a single network device, such as network device 700. In other embodiments, one or more microphones 802 are disposed on or within one network device, while the remaining microphones are disposed on or within one or more other network devices.

[0119] In practice, microphone 802 captures audio content arriving at microphone 802. As shown, when person 804 speaks near microphone 802, person 804 generates a speech signal s(t). As speech signal s(t) propagates throughout environment 800, at least some of speech signal s(t) reflects from walls or other nearby objects in environment 800. These reflections can distort speech signal s(t), causing the version of the speech signal captured by microphone 802 to be an echoed speech signal x(t), which is different from the original speech signal s(t).

[0120] In addition, environment 800 includes one or more noise sources 806, such as noise from nearby traffic or construction, noise from people moving throughout the environment, noise from one or more playback devices in environment 800, or any other ambient noise. In some embodiments, noise source 806 includes speech content from a person other than person 804. In any case, noise source 806 generates a noise signal v(t) that is captured by some or all of microphones 802. In this regard, the audio signal captured by microphone 802 is represented as y(t), which is the sum of the reverberant speech signal x(t) and the noise signal v(t). And for each individual microphone in microphone 802, the captured audio signal can therefore be characterized as:

[0121] y n (t) = x n (t)+v n (t), n = 1, 2, ..., N (Equation 1)

[0122] Where n is the index of the reference microphone and N is the total number of microphones. Converting from the time domain to the frequency domain, the above equation can be expressed as:

[0123] Y n (f) = X n (f)+V n(f), n = 1, 2, ..., N (Equation 2)

[0124] Or, in vector form:

[0125] Y(f)=X(f)+V(f). (Equation 3)

[0126] In addition, the power spectral density (PSD) matrix P is defined yy (f), P xx (f) and P vv (f), where P yy (f) is the PSD matrix of the total captured audio content, P xx (f) is the PSD matrix of the voice part of the total captured audio content, and P vv (f) is the PSD matrix of the noise portion of the total captured audio content. These PSD matrices are determined using the following equations:

[0127] P yy (f) = E{y(f)y H (f)}, (Equation 4)

[0128] P xx (f) = E{x(f)x H (f)}, (Equation 5)

[0129] P vv (f) = E{v(f)v H (f)} (Equation 6)

[0130] Where E{} denotes the expected value operator and H denotes the Hermitian transpose operator. Assuming that there is a lack of correlation between the voice portion and the noise portion of the total captured audio content (which is usually the case), the PSD matrix of the voice portion of the total captured audio content can be written as:

[0131] P xx (f) = P yy (f)P vv (f). (Equation 7)

[0132] To reduce the noise content V(f) and restore the voice content X(f) of the captured multi-channel audio content Y(f), the captured multi-channel audio content Y(f) is passed through a filter 808. In some embodiments, the filter 808 comprises a tangible, non-transitory computer-readable medium that, when executed by one or more processors of a network device, causes the network device to perform the multi-channel filtering functionality disclosed and described herein.

[0133] The filter 808 can filter the captured multi-channel audio content Y(f) in various ways. In some embodiments, the filter 808 is a linear filter h i (f) (where i = 1, 2, ..., N is the index of the reference microphone) is applied to the vector Y(f) of captured multi-channel audio content. In this way, N linear filters hi(f) (one for each microphone 802) are applied to the audio content vector Y(f). Applying these filters will produce the filtered output Z given below i (f):

[0134]

[0135] The filtered output Z i (f) including the filtered speech component D i (f) and the residual noise component v i (f), where

[0136]

[0137] and

[0138]

[0139] To determine the linear filter hi(f), an optimization constraint set is defined. In some embodiments, the optimization constraint is defined so as to maximize the degree of noise reduction while limiting the degree of signal distortion, for example, by limiting the degree of signal distortion to be less than or equal to a threshold degree. nr (h i (f)) is defined as:

[0140]

[0141] And the signal distortion index v Sd (h i (f)) is defined as:

[0142]

[0143] Among them, u i is the i-th standard basis vector and is defined as:

[0144]

[0145] Therefore, in order to maximize noise reduction while limiting signal distortion, the optimization problem in some implementations is to satisfy v sd (h i (f))≤σ 2 (f) maximizes ξ nr (hi (f)). To find the solution associated with this optimization problem, the derivatives of the associated Lagrangian with respect to hi(f) are set to zero, and the resulting closed-form solution is:

[0146] h i (f)=[P xx (f)+βP vv (f)] -1 P xx (f)u i (Equation 14)

[0147] where β (which is positive and is the inverse of the Lagrange multiplier) is the factor that allows tuning h i (f) Signal distortion and noise reduction factor at the output.

[0148] This linear filter h i The implementation of (f) can be computationally demanding. To reduce the filter h i The computational complexity of (f) is reduced by, in some embodiments, using the matrix P xx The fact that (f) is a rank 1 matrix leads to a more simplified form. And, since P xx (f) is a rank 1 matrix, so P -1 vv (f)P xx The rank of (f) is also 1. In addition, the matrix inversion can be further simplified using the Woodbury matrix identity. Applying all these concepts, the linear filter h i (f) can be expressed as:

[0149]

[0150] in

[0151]

[0152] It's P -1 vv (f)P xx The only positive eigenvalue of (f) is used as the normalization factor.

[0153] The linear filter h i One advantage of (f) is that it relies only on the PSD matrix of the total captured audio and the PSD matrix of the noise portion of the total captured audio, and therefore does not depend on the speech portion of the total captured audio. Another advantage is that parameter β allows customization of the degree of noise reduction and signal distortion. For example, increasing β increases noise reduction at the expense of increased signal distortion, while decreasing β reduces signal distortion at the expense of increased noise.

[0154] Since the linear filter h i (f) PSD matrix P depending on the total captured audio yy (f) and the PSD matrix P of the noise portion of the total captured audio vv (f), and thus estimate these PSD matrices to apply the filter. In some embodiments, first-order exponential smoothing is used to convert P yy Estimated to be:

[0155] P yy (n) = α y P yy (n-1)+(1-α y )yy H (Equation 17)

[0156] Among them, α y is the smoothing coefficient, and n represents the time frame index. In addition, the frequency index (f) has been deleted from this equation and the following equations to simplify the notation, but it should be understood that the process disclosed herein is performed for each frequency bin. The smoothing coefficient ranges from 0 to 1 and can be adjusted to tune the effect of P yy By reducing the P between consecutive time frame indices yy The degree of change, increase α y Can improve P yy The smoothness of the estimate is improved by increasing the P between consecutive time frame indices. yy The degree of change, reducing α y Can reduce P yy The estimated smoothness.

[0157] To estimate P vv In some embodiments, the filter 808 determines whether speech content exists in each frequency bin. If the filter 808 determines that speech content exists or may exist in a particular frequency bin, the filter 808 determines that the frequency bin does not represent noise content, and the filter 808 does not use the frequency bin to estimate P. vv On the other hand, if filter 808 determines that speech content is not present or is unlikely to be present in a particular frequency bin, filter 808 determines that the frequency bin consists primarily or entirely of noise content, and filter 808 uses the noise content to estimate P vv .

[0158] Filter 808 can determine whether speech content is present in a frequency bin in various ways. In some embodiments, filter 808 uses a hard voice activity detection (VAD) algorithm to make this determination. In other embodiments, filter 808 uses a soft speech presence probability algorithm to make this determination. For example, assuming a Gaussian distribution, the speech presence probability is calculated as follows:

[0159]

[0160] Where n is the timeframe index,

[0161]

[0162]

[0163] And among them

[0164]

[0165] is the a priori probability of speech absence. The derivation of this speech presence probability is described in Souden et al., “Gaussian Model-Based Multichannel Speech Presence Probability,” IEEE Transactions on Audio, Speech, and Language Processing (2010), which is incorporated herein by reference in its entirety.

[0166] It is worth noting that the calculation of the probability of speech presence depends on the PSD matrix P of the speech content. xx However, due to P xx (f) = P yy (f)-P vv (f), so this dependency can be eliminated by rewriting γ as follows:

[0167]

[0168] Furthermore, the variable ξ can be written as:

[0169]

[0170] in

[0171]

[0172] in

[0173]

[0174] And among them

[0175]

[0176] By defining the vector, the computational complexity of the voice presence probability calculation can be further reduced:

[0177]

[0178] So that ψ can be written as:

[0179]

[0180] And γ can be written as:

[0181]

[0182] Therefore, by computing y before attempting to compute ψ or γ temp , when the filter 808 determines the probability of speech existence, repeated calculations can be avoided.

[0183] Once the speech presence probability is determined for a given time frame, filter 808 updates the estimate of the noise covariance matrix by employing an expectation operator according to the following equation:

[0184]

[0185] in

[0186]

[0187] is the effective frequency-dependent smoothing coefficient.

[0188] In order to obtain the updated P -1 vv (n) for use in hi(f), the Sherman-Morrison formula is used as follows:

[0189]

[0190]

[0191] in

[0192]

[0193] Once the updated P is determined -1 vv (n), for all values ​​of f and all values ​​of i, the filter 808 may determine the linear filter h i (n) and applies it to the captured audio content. The output of filter 808 is then represented as y 0,i (n) = h H i (n)y(n). In some embodiments, the filter 808 uses the matrix H(n) to calculate the output in parallel for all i, where the columns are h i (n), so that

[0194]

[0195] And y out =HH y, (Equation 36)

[0196] in

[0197] And ξ=λ(n)-N. (Equation 38)

[0198] In some embodiments, filter 808 does not directly calculate H, which requires matrix multiplication. Instead, filter 808 calculates the output as follows, significantly reducing computational complexity:

[0199]

[0200] and

[0201] Using the above concepts, the filter 808 suppresses noise and preserves the speech content in the multi-channel audio signal captured by the microphone 802. In a simplified manner, this may include

[0202] A. Update P for all f yy (n)

[0203] B. Calculate the voice presence probability P(H1|y(n)) for all f

[0204] C. Update P for all f using the voice presence probability -1 vv (n)

[0205] D. Calculate the linear filter hi(n) for all f and all i and calculate the output as follows

[0206] yo, i(n) = h H i (n)y(n)

[0207] A more detailed example may include performing the following steps.

[0208] Step 1: Initialize the parameters and state variables at time frame 0. In some embodiments, the P is estimated over a specific time period (e.g., 500 ms). yy To initialize P yy and p -1 vv , and then use the estimated P yy P -1 vv Initialize to its reciprocal.

[0209] Step 2: At each time frame n, the following steps 3-13 are performed.

[0210] Step 3: For each frequency index f = {1, ..., K}, update the pair P according to Equation 17 yy (n) is estimated, and y is calculated according to Equation 27 temp , and calculate ψ according to Equation 28.

[0211] Step 4: For each frequency index f={1, . . . , K}, ψ is calculated using vector operations according to Equation 24.

[0212] Step 5: For each frequency index f={1, . . . , K}, ξ is calculated using vector operations according to Equation 23.

[0213] Step 6: For each frequency index f={1, . . . , K}, γ is calculated according to Equation 29.

[0214] Step 7: Vector operations are used according to Equation 18 to calculate the voice presence probability over all frequency bins.

[0215] Step 8: Calculate the value for updating P according to Equations 30 and 31 vv (n) Effective smoothing coefficient.

[0216] Step 9: Calculate w according to Equation 34.

[0217] Step 10: For each frequency index f={1, ..., K}, k(n) is updated according to Equation 32, and P is updated according to Equation 33. -1 vv (n).

[0218] Step 11: For each frequency index f={1, ..., K}, λ(n) is updated according to Equation 37.

[0219] Step 12: ξ is calculated according to Equation 38.

[0220] Step 13: For each frequency index f={1, ..., K}, the output y is calculated according to Equation 39 and according to Equation 40. out ,calculate Output vector of size N x 1.

[0221] In addition to the other advantages already described, the MCWF-based processing described above provides other advantages. For example, the captured audio signal is filtered in a distributed manner so that the audio signal does not need to be aggregated at a central node for processing. In addition, the MCWF algorithm can be executed at a separate node where a microphone is present, and then the node can share the output from the MCWF algorithm with some or all other nodes in the networked system. For example, Figure 7C Each of microphones 702 in FIG. 1 is part of a corresponding node capable of executing the MCWF algorithm. Thus, the node including microphone 702 a processes the audio captured by microphone 702 a according to the MCWF algorithm and then provides the MCWF output to the nodes associated with microphones 702 b-g. Similarly, the node including microphone 702 a receives the MCWF output from each of the nodes associated with microphones 702 b-g. Thus, each node can use the MCWF output from the other nodes when estimating and filtering noise content according to the MCWF algorithm.

[0222] Reference again Figure 8 Once the filter 808 suppresses the noise content and preserves the voice content in the individual audio signals captured by the microphone 802, for example using the MCWF algorithm described above, the filter 808 combines the filtered audio signals into a single signal. Having suppressed the noise content of each audio signal and preserved the voice content, the combined signal similarly has suppressed noise content and preserved voice content.

[0223] The filter 808 provides the combined signal to the speech processing block 810 for further processing. The speech processing block 810 runs a wake-up word detection process on the output of the filter 808 to determine whether the speech content output by the filter includes a wake-up word. In some embodiments, the speech processing block 810 is implemented as software executed by one or more processors of the network device 700. In other embodiments, the speech processing block 810 is a separate computing system, for example, Figure 5 One or more of computing devices 504, 506, and / or 508 are shown and described.

[0224] In response to determining that the output of filter 808 includes a wake word, speech processing block 810 performs further speech processing on the output of filter 808 to identify a voice command following the wake word. And in response to speech processing block 810 identifying the voice command following the wake word, network device 700 performs a task corresponding to the identified voice command. For example, as described above, in certain embodiments, network device 700 can send the voice input or a portion thereof to a remote computing device associated with, for example, a voice assistant service.

[0225] In some embodiments, the robustness and performance of the MCWF may be enhanced based on one or more of the following adjustments to the aforementioned algorithms.

[0226] 1) Parameter β can be time-frequency dependent. There are various methods for designing a time-frequency dependent β based on the probability of speech presence, signal diffusion ratio (SDR), and other factors. The idea is to use a smaller value when the SDR is high and speech is present to reduce speech distortion, and a larger value when the SDR is low or speech is absent to increase noise reduction. This value balances noise reduction and speech distortion based on the conditional probability of speech presence. A simple and effective approach is to define β as:

[0227] β(y)=β0 / (α β +(1-α β )β0 P(H1|y))

[0228] The conditional speech presence probability is combined to adjust the parameter β based on the input vector y. β A compromise is provided between a fixed tuning parameter and one that depends entirely on the probability of speech presence. β =0.5.

[0229] 2) The MMSE estimate of the desired speech signal can be obtained according to the following equation

[0230] y out =P(H1|y)H H (n)y(n)+(1-P(H1|y))G min y

[0231] Among them, the gain factor G min Determine the maximum amount of noise reduction when the probability of speech presence indicates that speech is not present. The importance of this model is that it can mitigate speech distortion in the event of an incorrect decision about the probability of speech presence. This approach improves robustness. This implementation can be done after step 13 of the algorithm and can be modified as follows out

[0232] y out =P(H1|y)y out +(1-P(H1|y))G min y

[0233] Among them, the probability of voice presence is used to generate the output and control how to apply G min .

[0234] 3) The algorithm is tuned and implemented in two supported modes: A) Noise Suppression (NS); B) Residual Echo Suppression (RES). If the speaker is playing content, the algorithm can operate in RES mode. Otherwise, the algorithm will operate in NS mode. The mode can be determined using internal state regarding the presence of audio playback.

[0235] 4) Initialize the covariance matrices in step 1 of the algorithm. The algorithm includes an initialization period, in which the input signal from the microphone array is used to estimate the initial input and noise covariance matrices. It can be assumed that no speech is present during this initialization period. These covariance matrices are initialized with diagonal matrices to simplify implementation. The initialization time can be adjusted in the algorithm to, for example, 0.5 seconds. This method provides a more robust solution that is insensitive to input level and noise type. As a result, relatively similar convergence rates can be achieved across all SNR levels and loudness levels.

[0236] 5) Considering the statistical characteristics of the voice signal, in order to improve the probability of multi-channel voice, the following recursive smoothing multi-channel voice probability can be used

[0237]

[0238] Among them, the smoothing coefficient α P The value of is between 0 and 1, and the smoothing coefficient can be adjusted to adjust the estimate of the probability of speech presence in the parameter adjustment stage.

[0239] V. Example Noise Suppression Methods

[0240] Figure 9 An example embodiment of method 900 is shown that may be implemented by a network device, such as network device 700 or any of the PBDs, NMDs, controller devices, or other VEDs disclosed and / or described herein, or any other voice-enabled device now known or later developed.

[0241] Various embodiments of method 900 include one or more operations, functions, and actions shown in blocks 902 through 914. Although these blocks are shown sequentially, these blocks may be performed in parallel and / or in a different order than that disclosed and described herein. Furthermore, various blocks may be combined into fewer blocks, divided into additional blocks, and / or removed based on the desired implementation.

[0242] In addition, for the method 900 and other processes and methods disclosed herein, the flowchart illustrates one possible implementation of the functions and operations of some embodiments. In this regard, each box may represent a module, segment, or portion of a program code, which includes one or more instructions executable by one or more processors for implementing a specific logical function or step in the process. The program code may be stored on any type of computer-readable medium, for example, a storage device including a disk or hard drive. Computer-readable media may include non-transitory computer-readable media, for example, tangible non-transitory computer-readable media for storing data for a short period of time, such as register memory, processor cache, and random access memory (RAM). Computer-readable media may also include non-transitory media, for example, auxiliary storage or persistent long-term storage devices, such as read-only memory (ROM), optical or magnetic disks, compact disk read-only memory (CD-ROM), etc. The computer-readable medium may also be any other volatile or non-volatile storage system. The computer-readable medium may be considered to be a computer-readable storage medium, such as a tangible storage device. In addition, for the method 800 and other processes and methods disclosed herein, Figure 9 Each block in the diagram may represent circuits connected to perform the specific logical functions in the process.

[0243] The method 900 begins at block 902, which includes a network device capturing (i) a first audio signal via a first microphone of a plurality of microphones, and (ii) a second audio signal via a second microphone of the plurality of microphones, wherein the first audio signal includes first noise content from a noise source and the second audio signal includes second noise content from the same noise source. In some embodiments, the plurality of microphones including the first microphone and the second microphone are the same network device (e.g., Figure 7A-7B Components of the network device 700a or 700b) depicted in FIG. In other embodiments, for example, Figure 7C As shown, at least some of the plurality of microphones are components of different network devices. In an example implementation, a first microphone is a component of a first network device (e.g., network device 700c), and a second microphone is a component of a second network device (e.g., network device 700d).

[0244] Next, method 900 proceeds to block 904, which includes identifying first noise content in the first audio signal. In some embodiments, identifying the first noise content in the first audio signal involves one or more of: (i) the network device using a VAD algorithm to detect the absence of speech in the first audio signal, or (ii) the network device using a speech presence probability algorithm to determine a probability that speech is present in the first audio signal. An example of a speech presence probability algorithm is described above with reference to Equation 18. If the VAD algorithm detects the absence of speech in the first audio signal, or the speech presence probability algorithm indicates that the probability of speech being present in the first audio signal is below a threshold probability, this may indicate that the first audio signal is noise-dominated and includes little or no speech content.

[0245] Next, method 900 proceeds to block 906, which includes determining an estimated noise content captured by the plurality of microphones using the identified first noise content. In some embodiments, determining an estimated noise content captured by the plurality of microphones using the identified first noise content involves the network device updating a noise content PSD matrix for use in the MCWF algorithm described above with reference to equations 30-34.

[0246] In some embodiments, the following steps are performed based on the probability of speech being present in the first audio signal being below a threshold probability: identifying first noise content in the first audio signal at block 904, and using the identified first noise content to determine estimated noise content captured by the multiple microphones at block 906. As described above, a speech presence probability algorithm indicating that the probability of speech being present in the first audio signal is below a threshold probability indicates that the first audio signal is noise-dominated and includes little or no speech content. Such a noise-dominated signal is more likely than a non-noise-dominated signal to provide an accurate estimate of the noise present in other signals captured by the microphones (e.g., the second audio signal). Therefore, in some embodiments, the step of determining estimated noise content captured by the multiple microphones using the identified first noise content is performed in response to determining that the probability of speech being present in the first audio signal is below a threshold probability. The threshold probability can take on various values, and in some embodiments, the threshold probability can be adjusted to tune the noise filtering methods described herein. In some embodiments, the threshold probability is set as low as 1%. In other embodiments, the threshold probability is set to a higher value, such as between 1% and 10%.

[0247] Next, method 900 proceeds to block 908, which includes suppressing first noise content in the first audio signal and second noise content in the second audio signal using the estimated noise content. In some embodiments, suppressing the first noise content in the first audio signal and the second noise content in the second audio signal using the estimated noise content involves the network device applying a linear filter to each of the audio signals captured by the plurality of microphones as described above with reference to equations 35-40 using the updated noise content PSD matrix.

[0248] Next, method 900 proceeds to block 910, which includes combining the suppressed first audio signal and the suppressed second audio signal into a third audio signal. In some embodiments, combining the suppressed first audio signal and the suppressed second audio signal into the third audio signal involves the network device combining the suppressed audio signals from all microphones in the plurality of microphones into the third audio signal.

[0249] Next, method 900 proceeds to block 912, which includes determining that the third audio signal includes voice input including a wake-up word. In some embodiments, the step of determining that the third audio signal includes voice input including a wake-up word involves the network device executing one or more voice processing algorithms on the third audio signal to determine whether any portion of the third audio signal includes the wake-up word. In operation, the step of determining that the third audio signal includes voice input including a wake-up word can be performed according to any wake-up word detection method disclosed and described herein and / or any wake-up word detection method now known or later developed.

[0250] Finally, method 900 proceeds to box 914, which includes: in response to determining that the third audio signal includes speech content including a wake-up word, sending at least a portion of the voice input to a remote computing device for speech processing to identify a voice utterance different from the wake-up word. As described above, the voice input can include the wake-up word and the voice utterance following the wake-up word. The voice utterance can include a spoken command and one or more spoken keywords. Therefore, in some embodiments, the step of sending at least a portion of the voice input to a remote computing device for speech processing to identify a voice utterance different from the wake-up word includes sending a portion of the voice input following the wake-up word (which can include a spoken command and / or spoken keywords) to a separate computing system for speech analysis.

[0251] VII. Conclusion

[0252] The above description discloses, among other things, various example systems, methods, apparatuses, and articles of manufacture, including, inter alia, firmware and / or software executed on hardware. It should be understood that these examples are merely illustrative and should not be considered restrictive. For example, it is contemplated that any or all of these firmware, hardware, and / or software aspects or components may be implemented exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Therefore, the examples provided are not the only way to implement these systems, methods, apparatuses, and / or articles of manufacture.

[0253] (Feature 1) A network device comprising: (i) a plurality of microphones, including a first microphone and a second microphone; (ii) one or more processors; and (iii) a tangible, non-transitory computer-readable medium storing instructions executable by the one or more processors to cause the network device to perform operations comprising: (a) capturing (i) a first audio signal via the first microphone, and (ii) a second audio signal via the second microphone, wherein the first audio signal includes first noise content from a noise source and the second audio signal includes second noise content from the noise source; (b) identifying first noise content in the first audio signal; (c) using the identified first noise content to determine an estimated noise content captured by the plurality of microphones; (d) using the estimated noise content to suppress the first noise content in the first audio signal and the second noise content in the second audio signal; (e) combining the suppressed first audio signal and the suppressed second audio signal into a third audio signal; (f) determining that the third audio signal includes speech input including a wake-up word; and (g) in response to the determination, sending at least a portion of the speech input to a remote computing device for speech processing to identify a speech utterance different from the wake-up word.

[0254] (Feature 2) A network device according to Feature 1, wherein the operation also includes: (i) determining a probability that the first audio signal includes speech content, (ii) wherein, based on the determined probability being lower than a threshold probability, performing the following steps: (a) identifying first noise content in the first audio signal, and (b) using the identified first noise content to determine an estimated noise content captured by the multiple microphones.

[0255] (Feature 3) The network device according to Feature 1 further includes a housing that at least partially encloses the components of the network device within the housing, wherein the first microphone and the second microphone are arranged along the housing and separated from each other by a distance greater than about five centimeters.

[0256] (Feature 4) According to the network device of Feature 1, the operation also includes: (i) capturing a fourth audio signal via a third microphone among the multiple microphones, wherein the fourth audio signal includes third noise content from the noise source; (ii) identifying the third noise content in the fourth audio signal; and (iii) using the identified third noise content to update the estimated noise content captured by the multiple microphones.

[0257] (Feature 5) The network device according to Feature 4, wherein the network device captures the fourth audio signal simultaneously with capturing the first audio signal and the second audio signal.

[0258] (Feature 6) The network device according to Feature 4 further includes a housing that at least partially encloses the components of the network device within the housing, wherein the first microphone, the second microphone, and the third microphone are arranged along the housing and separated from each other by a distance greater than about five centimeters.

[0259] (Feature 7) A network device according to Feature 1, wherein sending at least a portion of the voice input to a remote computing device for voice processing to identify a voice utterance different from the wake-up word includes: sending a portion of the voice input after the wake-up word to a separate computing system for voice analysis.

[0260] (Feature 8) A tangible, non-transitory computer-readable medium storing instructions executable by one or more processors to cause a network device to perform operations comprising: (i) capturing, via multiple microphones of the network device, (a) a first audio signal via a first microphone of the multiple microphones, and (b) a second audio signal via a second microphone of the multiple microphones, wherein the first audio signal includes first noise content from a noise source and the second audio signal includes second noise content from the noise source; (ii) identifying the first noise content in the first audio signal; (iii) using the identified first noise content to determine an estimated noise content captured by the multiple microphones; (iv) using the estimated noise content to suppress the first noise content in the first audio signal and the second noise content in the second audio signal; (v) combining the suppressed first audio signal and the suppressed second audio signal into a third audio signal; (vi) determining that the third audio signal includes speech input, the speech input including a wake-up word; and (vii) in response to the determination, sending at least a portion of the speech input to a remote computing device for speech processing to identify a speech utterance different from the wake-up word.

[0261] (Feature 9) According to the tangible, non-transitory computer-readable medium of Feature 8, the operation also includes: (i) determining a probability that the first audio signal includes speech content, (ii) wherein, based on the determined probability being lower than a threshold probability, performing the following steps: (a) identifying first noise content in the first audio signal, and (b) using the identified first noise content to determine an estimated noise content captured by the multiple microphones.

[0262] (Feature 10) The tangible, non-transitory computer-readable medium of Feature 8, wherein the network device includes a housing that at least partially encloses components of the network device within the housing, and wherein the first microphone and the second microphone are arranged along the housing and separated from each other by a distance greater than about five centimeters.

[0263] (Feature 11) According to the tangible, non-transitory computer-readable medium of Feature 8, the operation also includes: (i) capturing a fourth audio signal via a third microphone among the multiple microphones, wherein the fourth audio signal includes third noise content from the noise source; (ii) identifying the third noise content in the fourth audio signal; and (iii) using the identified third noise content to update the estimated noise content captured by the multiple microphones.

[0264] (Feature 12) The tangible, non-transitory computer-readable medium of Feature 11, wherein the fourth audio signal is captured simultaneously with the first audio signal and the second audio signal.

[0265] (Feature 13) The tangible, non-transitory computer-readable medium of Feature 11, wherein the network device includes a housing that at least partially encloses components of the network device within the housing, wherein the first microphone, the second microphone, and the third microphone are arranged along the housing and separated from each other by a distance greater than about five centimeters.

[0266] (Feature 14) The tangible, non-transitory computer-readable medium of Feature 8, wherein sending at least a portion of the voice input to a remote computing device for voice processing to identify a voice utterance different from the wake-up word includes sending a portion of the voice input following the wake-up word to a separate computing system for voice analysis.

[0267] (Feature 15) A method comprising: (i) capturing, via multiple microphones of a network device, (a) a first audio signal via a first microphone of the multiple microphones, and (b) a second audio signal via a second microphone of the multiple microphones, wherein the first audio signal includes first noise content from a noise source and the second audio signal includes second noise content from the noise source; (ii) identifying the first noise content in the first audio signal; (iii) using the identified first noise content to determine an estimated noise content captured by the multiple microphones; (iv) using the estimated noise content to suppress the first noise content in the first audio signal and the second noise content in the second audio signal; (v) combining the suppressed first audio signal and the suppressed second audio signal into a third audio signal; (vi) determining that the third audio signal includes voice input, the voice input including a wake-up word; and (vii) in response to the determination, sending at least a portion of the voice input to a remote computing device for voice processing to identify a voice utterance different from the wake-up word.

[0268] (Feature 16) The method according to Feature 15 also includes: (i) determining a probability that the first audio signal includes speech content, (ii) wherein, based on the determined probability being lower than a threshold probability, performing the following steps: (a) identifying first noise content in the first audio signal, and (b) using the identified first noise content to determine an estimated noise content captured by the multiple microphones.

[0269] (Feature 17) A method according to Feature 15, wherein the network device includes a housing that at least partially encloses components of the network device within the housing, and wherein the first microphone and the second microphone are arranged along the housing and separated from each other by a distance greater than about five centimeters.

[0270] (Feature 18) The method according to Feature 15 also includes: (i) capturing a fourth audio signal via a third microphone among the multiple microphones, wherein the fourth audio signal includes third noise content from the noise source; (ii) identifying the third noise content in the fourth audio signal; and (iii) using the identified third noise content to update the estimated noise content captured by the multiple microphones.

[0271] (Feature 19) The method according to Feature 18, wherein the fourth audio signal is captured simultaneously with the first audio signal and the second audio signal.

[0272] (Feature 20) The method of Feature 18, wherein the network device includes a housing that at least partially encloses components of the network device within the housing, wherein the first microphone, the second microphone, and the third microphone are arranged along the housing and separated from each other by a distance greater than about five centimeters.

[0273] Furthermore, references herein to an "embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one exemplary embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor do they refer to separate or alternative embodiments that are mutually exclusive of other embodiments. Therefore, it should be understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0274] This specification is presented primarily in terms of illustrative environments, systems, processes, steps, logic blocks, processing, and other symbolic representations that are directly or indirectly analogous to the operations of a data processing device coupled to a network. These process descriptions and representations are typically used by those skilled in the art to communicate the content of their work to others skilled in the art. Various specific details are set forth to provide a thorough understanding of the present disclosure. However, it should be understood by those skilled in the art that specific, specific details are not required to practice the present disclosure. In other instances, well-known methods, processes, components, and circuits are not described to avoid unnecessarily obscuring aspects of the embodiments. For example, in some embodiments, other techniques may be used to determine the probability of loss of speech. Accordingly, the scope of the present disclosure is defined by the appended claims, rather than the description of the above embodiments.

[0275] When any of the appended claims is understood to cover a pure software and / or firmware implementation, at least one element in at least one example is hereby expressly defined to include a non-transitory tangible medium storing the software and / or firmware, such as a memory, DVD, CD, Blu-ray, etc.

Claims

1. A method (900) for a network device (700), the network device comprising a plurality of microphones, the plurality of microphones comprising at least a first microphone and a second microphone, the method comprising: capturing a plurality of corresponding audio signals via at least a portion of the plurality of microphones (902), wherein the plurality of corresponding audio signals includes a first audio signal captured via a first microphone and a second audio signal captured via a second microphone; identifying first noise content in a first audio signal captured via a first microphone (904); determining an estimated noise content captured by the plurality of microphones based on the identified first noise content (906); generating a first noise suppressed audio signal by suppressing first noise content in the first audio signal using the estimated noise content (908); generating a second noise suppressed audio signal by suppressing second noise content in the second audio signal using the estimated noise content (908); generating a third audio signal by combining the first noise-suppressed audio signal with the second noise-suppressed audio signal (910); determining whether the third audio signal includes a voice input including a wake word ( 912 ); and When it is determined that the third audio signal includes speech input including the wake word, at least a portion of the speech input is sent to a remote computing device for speech processing to identify a speech utterance different from the wake word (914).

2. The method according to claim 1, wherein Identifying first noise content in the first audio signal captured via the first microphone device includes: determining a probability that the first audio signal includes speech content; and identifying the first audio signal as capable of determining the estimated noise content when the determined probability is below a threshold probability.

3. The method according to claim 2, wherein: Determining the estimated noise content captured by the plurality of microphones includes updating a noise content power spectral density matrix for use with a multi-channel Wiener filter.

4. The method according to any one of the preceding claims, further comprising: capturing, via a third microphone (702c) of the plurality of microphones, a fourth audio signal including third noise content; identifying the third noise content in the fourth audio signal; as well as The identified third noise content is used to update the estimated noise content captured by the plurality of microphones.

5. The method of any preceding claim, further comprising simultaneously capturing a plurality of audio signals via respective microphones of the plurality of microphones.

6. A method according to any preceding claim, wherein: Generating a first noise suppressed audio signal by suppressing the first noise content in the first audio signal using the estimated noise content comprises: passing the first audio signal through at least one linear filter h i (f), and generating a second noise-suppressed audio signal by suppressing the second noise content in the second audio signal using the estimated noise content comprises: passing the second audio signal through the at least one linear filter h i (f).

7. The method according to claim 6, wherein: Passing the first noise-suppressed audio signal and the second noise-suppressed audio signal through the at least one linear filter h i (f) comprising: passing the updated noise content power spectral density matrix through the at least one linear filter h i (f).

8. The method according to claim 6, wherein: The at least one linear filter h i (f) Depends on: a power spectral density matrix of the captured first audio signal, and a power spectral density matrix of the noise component of the captured first audio signal, and is independent of the speech portion of the captured first audio signal.

9. The method according to claim 6, wherein: The at least one linear filter h i (f) is expressed as: in: (Equation 1) in: (Equation 2) (f) is the frequency index; N is the index of the corresponding microphone; P yy (f) is a power spectral density matrix of the audio content captured by the corresponding microphone; P vv (f) is a power spectral density matrix of the noise portion of the audio content captured by the corresponding microphone; and β is the inverse of the Lagrange multiplier factor and can be modified to tune signal distortion and noise reduction; and u i is the i-th standard basis vector, where: .

10. The method according to any one of claims 5 to 9, further comprising: First-order exponential smoothing is applied to the captured first audio signal.

11. The method according to any one of claims 2 to 10, wherein: Determining a probability that the first audio signal includes speech content comprises using at least one of: a voice activity detection algorithm for detecting the presence of speech in the first audio signal; and A speech presence probability algorithm is used to determine the probability of speech being present in the first audio signal.

12. The method according to any preceding claim, further comprising: Initialization is performed at the initial time frame of the predetermined period, Estimate P at the initial time frame of the predetermined time period yy (f), then Using the estimated P yy Initialize P by the inverse of (f) -1 vv (f), in: P vv (f) is the power spectral density matrix of the noise portion of the audio content captured by the corresponding microphone; and P yy (f) is the power spectral density matrix of the audio content captured by the corresponding microphone.

13. The method according to claim 12, wherein: The predetermined time period is 500 ms.

14. The method according to claim 12 or 13, wherein: Determining an estimated noise content captured by the plurality of microphones comprises: Update P for all frequency segments yy (n); Calculate the probability of voice presence in all frequency segments; When the calculated voice presence probability is lower than the threshold, the P values ​​of all frequency segments are updated. -1 vv (n); and Calculate the at least one linear filter h for all frequency bins i (n), where: P yy is a power spectral density matrix of the audio content captured by the corresponding microphone, wherein the frequency indices (f) have been removed to simplify notation; n is the timeframe index; P -1 vv is to use the estimated P yy where the frequency index (f) has been removed to simplify the notation.

15. A method according to any preceding claim, wherein: Sending at least a portion of the voice input to a remote computing device includes sending a portion of the voice input following the wake word to a separate computing system for voice analysis.

16. A method according to any preceding claim, in, The plurality of microphones are arranged along a housing (704) of the network device (700) and are separated from each other by a distance greater than five centimeters, the housing at least partially enclosing components of the network device (700).

17. A tangible, non-transitory computer-readable medium storing instructions executable by one or more processors to cause a network device to perform the method of any preceding claim.

18. A network device (700) comprising a housing (704) having a plurality of microphones disposed thereon, the network device being configured to perform the method of any preceding claim.

Citation Information

Patent Citations

  • Media Playback System with Voice Assistance

    US20190102145A1

  • System and method for synchronizing operations among a plurality of independently clocked digital data processing devices

    US8234395B2

  • Processing Speech from Distributed Microphones

    US20170332168A1