Techniques for speech enhancement

DNN-DE techniques enhance speech clarity in mixed audio environments by selectively applying audio processing, addressing the challenge of dialogue clarity in media playback systems.

WO2026072097A1PCT designated stage Publication Date: 2026-04-02SONOS INC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing media playback systems struggle to provide clear and intelligible speech audio in mixed audio environments, particularly in home theater settings, where non-speech audio can overpower dialogue, leading to user frustration and the need for frequent volume adjustments.

Method used

Employing Deep Neural Net-based Dialogue Extraction (DNN-DE) techniques for real-time speech enhancement, selectively applying audio processing to enhance dialogue clarity while maintaining the original soundtrack context, and allowing user control over processing settings.

Benefits of technology

Improves dialogue clarity and maintains the overall sound quality by selectively enhancing speech audio, preserving the intended audio mix and user preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025024937_02042026_PF_FP_ABST
    Figure US2025024937_02042026_PF_FP_ABST
Patent Text Reader

Abstract

An example method includes detecting, with a playback device, an audio signal, applying a parametric machine learning model to dynamically detect speech in the audio signal, based on detecting the speech, separating the audio signal into speech audio and non-speech audio, applying first audio processing to the speech audio to produce processed speech audio, applying second audio processing to the non-speech audio to produced processed non-speech audio, the second audio processing being different from the first audio processing, combining the processed speech audio and the processed non-speech audio to produce an audio output signal, and playing back the audio output signal via the playback device.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCTTECHNIQUES FOR SPEECH ENHANCEMENTCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to co-pending U.S. Provisional Application No. 63 / 700,280 titled “TECHNIQUES FOR SPEECH ENHANCEMENT” and filed on September 27, 2024, which is hereby incorporated herein by reference in its entirety.FIELD OF THE DISCLOSURE

[0002] The present disclosure is related to consumer goods and, more particularly, to methods, systems, products, aspects, services, and other elements directed to media playback or some aspect thereof.BACKGROUND

[0003] Media playback systems, such as the SONOS Wireless Home Sound System, enable people to experience music from many sources via one or more networked playback devices. Through a software control application installed on a controller (e.g., smartphone, tablet, computer, voice input device, etc.), one can play what she wants in any room having a networked playback device. Media content (e.g., songs, podcasts, video sound, etc.) can be streamed to playback devices such that each room with a playback device can play back corresponding different media content. In addition, rooms can be grouped together for synchronous playback of the same media content, and / or the same media content can be heard in all rooms synchronously.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Aspects, and advantages of the presently disclosed technology may be better understood with regard to the following description, appended claims, and accompanying drawings, as listed below. A person skilled in the relevant art will understand that the elements shown in the drawings are for purposes of illustrations, and variations, including different and / or additional elements and arrangements thereof, are possible.

[0005] FIG. 1 A is a partial cutaway view of an environment having a media playback system configured in accordance with aspects of the disclosed technology.

[0006] FIG. IB is a schematic diagram of the media playback system of FIG. 1A and one or more networks according to aspects of the disclosed technology.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0007] FIG. 1C is a block diagram of a playback device according to aspects of the disclosed technology.

[0008] FIG. ID is a block diagram of a playback device according to aspects of the disclosed technology.

[0009] FIG. IE is a block diagram of a bonded playback device according to aspects of the disclosed technology.

[0010] FIG. IF is a block diagram of a network microphone device according to aspects of the disclosed technology.

[0011] FIG. 1G is a block diagram of a playback device according to aspects of the disclosed technology.

[0012] FIG. 1H is a partial schematic diagram of a control device according to aspects of the disclosed technology.

[0013] FIGS. II through IL are schematic diagrams of corresponding media playback system zones according to aspects of the disclosed technology.

[0014] FIG. IM is a schematic diagram of media playback system areas according to aspects of the disclosed technology.

[0015] FIG. 2 is a diagram illustrating an audio processing chain according to aspects of the disclosed technology.

[0016] FIG. 3 is a block diagram of one example of a machine learning personalization system according to aspects of the disclosed technology.

[0017] FIG. 4A is a block diagram of one example of a playback device according to aspects of the disclosed technology.

[0018] FIG. 4B is a front isometric view of an example of a portable playback device according to aspects of the disclosed technology.

[0019] FIG. 4C is a front isometric view of another example of a portable playback device according to aspects of the disclosed technology.

[0020] FIG. 4D is a front isometric view of an example of a portable playback device according to aspects of the disclosed technology.

[0021] FIG. 5 is a flow diagram of one example of a speech enhancement process according to aspects of the disclosed technology.

[0022] FIG. 6 is a flow diagram of another example of a speech enhancement process according to aspects of the disclosed technology.

[0023] FIG. 7 is a flow diagram of another example of a speech enhancement process according to aspects of the disclosed technology.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0024] FIG. 8 is a flow diagram of another example of a speech enhancement process according to aspects of the disclosed technology.

[0025] The drawings are for the purpose of illustrating example embodiments, but those of ordinary skill in the art will understand that the technology disclosed herein is not limited to the arrangements and / or instrumentality shown in the drawings.DETAILED DESCRIPTIONI. Overview

[0026] Examples described herein relate to techniques for enhancing the audio quality and intelligibility of speech audio. Good dialogue clarity is an important aspect of video sound (e.g., television sound) for many users. However, in many listening scenarios, users can be frustrated by the lack of clarity of speech audio relative to non-speech audio. For example, in many action movies and / or television series, special effects sounds and other non-speech sounds can be output at a much higher volume than speech audio, causing users to struggle to hear the dialogue and / or having to frequently change the playback volume as the video transitions between scenes with loud non-speech audio and scenes dominated by dialogue. Numerous factors complicate efforts to provide clear, intelligible speech audio in mixed audio settings (e.g., where the output sound at a given time includes both speech audio and non- speech audio), including factors such as the playback environment, listener preferences and characteristics (e.g., some listeners may have impaired hearing), and the audio content and / or structure of the soundtrack. Furthermore, maintaining natural sounding speech that is consistent with the video context can introduce additional complications. For example, if a video scene depicts a speaker in a noisy environment (e.g., a crowded restaurant or train station), audio processing that removes much of the background sound to enhance the clarity of the dialogue can result in the audio sounding out of context with the video. Thus, there are significant challenges to providing high-clarity speech audio, while maintaining the original context / intent of the soundtrack, particularly where the soundtrack is associated with video (e.g., in a home theater environment).

[0027] Accordingly, examples provide techniques for speech enhancement in mixed audio environments. As described further below, certain examples apply machine learning-based audio processing, such as Deep Neural Net-based Dialogue Extraction (DNN-DE) techniques, to provide high performance speech enhancement in real-time. Employing DNN-DE techniques according to examples disclosed herein may achieve improved dialogue clarity while also offering additional benefits, such as the ability to uncover dialogue that is maskedAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT by other sounds and the ability to apply speech enhancement processing selectively, rather than all the time. For example, using DDE-DE, speech audio in a soundtrack (or audio stream) can be detected, and based on detecting the speech, speech enhancement audio processing can be applied. However, when no speech is detected, speech enhancement processing may not be applied, thereby retaining the original audio mix and intended sound experience offered by the soundtrack. In this manner, dialogue can be made more intelligible without altering other aspects of the audio, as described further below.

[0028] In some playback environments, such as a home theater environment, for example, the incoming audio stream includes multi-channel audio content. For example, some home theater arrangements may include a center audio channel, front left and right audio channels, and one or more surround audio channels. In some examples, speech enhancement audio processing can be applied to the center channel, where most dialogue generally resides, without altering characteristics of other audio channels. As described further below, in some examples, audio processing of the center audio channel, including speech detection and optional speech enhancement processing, can be used to adjust the audio processing for one or more other audio channels to compensate for certain sound characteristics that may be lost in the center channel due to the speech enhancement processing. In this manner, the overall sound quality and experience can be maintained while dialogue clarity is improved.

[0029] According to certain examples, mechanisms for user control over at least some aspects of the audio processing are provided to improve both the usability and sound experience of speech enhancement approaches. For example, user of a playback device or media playback system may control whether to apply speech enhancement techniques at all, when to apply speech enhancement techniques (e.g., in certain playback environments, but not others), and / or what type of speech enhancement techniques to apply (e.g., DDE-DE, center channel processing only, etc.). Furthermore, examples provide for using passive user feedback to tailor audio processing based on identified user preferences, as described in more detail below.

[0030] According to certain aspects, for example, a method comprises detecting an audio signal with a playback device, applying a parametric machine learning model to dynamically detect speech in the audio signal, and based on detecting the speech, separating the audio signal into speech audio and non-speech audio. Examples of the method further comprise applying first audio processing to the speech audio to produce processed speech audio, applying second audio processing to the non-speech audio to produced processed non-speech audio, the second audio processing being different from the first audio processing, combining the processed speech audio and the processed non-speech audio to produce an audio output signal, andAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT playing back the audio output signal via the playback device. In some instances, detecting the audio signal comprising multi-channel audio content including a center audio channel and a plurality of additional audio channels, and applying the parametric machine learning model includes applying the parametric machine learning mode to the center audio channel to detect speech in the center audio channel. Examples of the method further comprise, based on detecting the speech, controlling audio processing applied to the center audio channel and to at least one additional audio channel of the plurality of additional audio channels.

[0031] While some examples described herein may refer to functions performed by given actors such as “users,” “listeners,” and / or other entities, it should be understood that such references are for purposes of explanation only. The claims should not be interpreted to require action by any such example actor unless explicitly required by the language of the claims themselves.

[0032] In the Figures, identical reference numbers identify generally similar, and / or identical, elements. Many of the details, dimensions, angles, and other aspects shown in the Figures are merely illustrative of particular embodiments of the disclosed technology. Accordingly, other embodiments can have other details, dimensions, angles, and aspects without departing from the spirit or scope of the disclosure. In addition, those of ordinary skill in the art will appreciate that further embodiments of the various disclosed technologies can be practiced without several of the details described below.II. Suitable Operating Environment

[0033] FIG. 1A is a partial cutaway view of a media playback system (MPS) 100 distributed in an environment 101 (e.g., a house). In the illustrated embodiment of FIG. 1A, the environment 101 comprises a household having several rooms, spaces, and / or playback zones, including (clockwise from upper left) a master bathroom 101a, a master bedroom 101b, a second bedroom 101c, a family room or den 101 d, an office lOle, a living room 10 If, a dining room 101g, a kitchen lOlh, and an outdoor patio lOli. While certain embodiments and examples are described below in the context of a home environment, the technologies described herein may be implemented in other types of environments. In some embodiments, for example, the media playback system 100 can be implemented in one or more commercial settings (e.g., a restaurant, mall, airport, hotel, a retail or other store), one or more vehicles (e.g., a sports utility vehicle, bus, car, a ship, a boat, an airplane, etc.), multiple environments (e.g., a combination of home and vehicle environments), and / or another suitable environment where multi-zone audio may be desirable.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0034] Within the rooms and spaces of the environment 101, the MPS 100 comprises one or more playback devices 110 (identified individually as playback devices 1 lOa-n), one or more network microphone devices 120 (“NMDs”) (identified individually as NMDs 120a-c), and one or more control devices 130 (identified individually as control devices 130a and 130b).

[0035] As used herein the term “playback device” can generally refer to a network device configured to receive, process, and output data of a media playback system. For example, a playback device can be a network device that receives and processes audio content. In some embodiments, a playback device includes one or more transducers or speakers powered by one or more amplifiers. In other embodiments, however, a playback device includes one of (or neither of) the speaker and the amplifier. For instance, a playback device can comprise one or more amplifiers configured to drive one or more speakers external to the playback device via a corresponding wire or cable.

[0036] Moreover, as used herein the term “NMD” (i.e., a “network microphone device”) can generally refer to a network device that is configured for audio detection. In some embodiments, an NMD is a stand-alone device configured primarily for audio detection. A stand-alone NMD 120 may omit components and / or functionality that is typically included in a playback device 110, such as a speaker or related electronics. For instance, in such cases, a stand-alone NMD may not produce audio output or may produce limited audio output. In other embodiments, an NMD is incorporated into a playback device (or vice versa). A playback device 110 that includes components and functionality of an NMD 120 may be referred to as being “NMD-equipped.” Examples of playback devices 110 and NMDs 120 are described further below.

[0037] The term “control device” can generally refer to a network device configured to perform functions relevant to facilitating user access, control, and / or configuration of the media playback system 100. Examples of control devices are described further below.

[0038] In some examples, one or more of the various playback devices 110 may be configured as portable playback devices, while others may be configured as stationary playback devices. For example, certain playback devices 110 may include an internal power source (e.g., a rechargeable battery) that allows the playback device to operate without being physically connected to a mains electrical outlet or the like. In this regard, such a playback device may be referred to herein as a “portable playback device.” On the other hand, playback devices that are configured to rely on power from a mains electrical outlet or the like may be referred to herein as “stationary playback devices,” although such devices may in fact be moved around a homeAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT or other environment. In practice, a person might often take a portable playback device to and from a home or other environment in which one or more stationary playback devices remain.

[0039] Each of the playback devices 110 is configured to receive audio signals or data from one or more media sources (e.g., one or more remote servers, one or more local devices, etc.) and play back the received audio signals or data as sound. The one or more NMDs 120 are configured to receive spoken word commands, and the one or more control devices 130 are configured to receive user input. In response to the received spoken word commands and / or user input, the media playback system 100 can play back audio via one or more of the playback devices 110. In certain embodiments, the playback devices 110 are configured to commence playback of media content in response to a trigger. For instance, one or more of the playback devices 110 can be configured to play back a morning playlist upon detection of an associated trigger condition (e.g., presence of a user in a kitchen, detection of a coffee machine operation, etc.). In some embodiments, for example, the media playback system 100 is configured to play back audio from a first playback device (e.g., the playback device 110a) in synchrony with a second playback device (e.g., the playback device 110b). Interactions between the playback devices 110, NMDs 120, and / or control devices 130 of the media playback system 100 configured in accordance with the various embodiments of the disclosure are described in greater detail below with respect to FIGS. 1B-1M.

[0040] The media playback system 100 can comprise one or more playback zones, some of which may correspond to the rooms in the environment 101. The media playback system 100 can be established with one or more playback zones, after which additional zones may be added, or removed, to form, for example, the configuration shown in FIG. 1 A. Each zone may be given a name according to a different room or space such as the office lOle, master bathroom 101a, master bedroom 101b, the second bedroom 101c, kitchen lOlh, dining room 101g, living room 10 If, and / or the balcony lOli. In some aspects, a single playback zone may include multiple rooms or spaces. In certain aspects, a single room or space may include multiple playback zones.

[0041] In the illustrated embodiment of FIG. 1A, the second bedroom 101c, the office lOle, the living room 10 If, the dining room 101g, the kitchen lOlh, and the outdoor patio lOli each include one playback device 110, and the master bathroom 101a, the master bedroom 101b, and the den 101 d include a plurality of playback devices 110. In the master bedroom 101b, the playback devices 1101 and 110m may be configured, for example, to play back audio content in synchrony as individual ones of playback devices 110, as a bonded playback zone, as a consolidated playback device, and / or any combination thereof. Similarly, in the den 101 d, theAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT playback devices HOh-k can be configured, for instance, to play back audio content in synchrony as individual ones of playback devices 110, as one or more bonded playback devices, and / or as one or more consolidated playback devices. Additional details regarding bonded and consolidated playback devices are described below with respect to FIGS. IB, IE, and 1I-M.

[0042] In some aspects, one or more of the playback zones in the environment 101 may each be playing different audio content. For instance, a user may be grilling on the patio lOli and listening to hip hop music being played by the playback device 110c while another user is preparing food in the kitchen lOlh and listening to classical music played by the playback device 110b. In another example, a playback zone may play the same audio content in synchrony with another playback zone. For instance, the user may be in the office lOle listening to the playback device 1 lOf playing back the same hip hop music being played back by playback device 110c on the patio lOli. In some aspects, the playback devices 110c and 11 Of play back the hip hop music in synchrony such that the user perceives that the audio content is being played seamlessly (or at least substantially seamlessly) while moving between different playback zones. Additional details regarding audio playback synchronization among playback devices and / or zones can be found, for example, in U.S. Patent No. 8,234,395 titled, “System and method for synchronizing operations among a plurality of independently clocked digital data processing devices,” which is incorporated herein by reference in its entirety. a. Suitable Media Playback System

[0043] FIG. IB is a schematic diagram of the media playback system 100 and a cloud network 102. For ease of illustration, certain devices of the media playback system 100 and the cloud network 102 are omitted from FIG. IB. One or more communication links 103 (referred to hereinafter as “the links 103”) communicatively couple the media playback system 100 and the cloud network 102.

[0044] The links 103 can comprise, for example, one or more wired networks, one or more wireless networks, one or more wide area networks (WAN), one or more local area networks (LAN), one or more personal area networks (PAN), one or more telecommunication networks (e.g., one or more Global System for Mobiles (GSM) networks, Code Division Multiple Access (CDMA) networks, Long-Term Evolution (LTE) networks, 5G communication networks, and / or other suitable data transmission protocol networks), etc. The cloud network 102 is configured to deliver media content (e.g., audio content, video content, photographs, social media content, etc.) to the media playback system 100 in response to a request transmitted from the media playback system 100 via the links 103. In some embodiments, the cloud networkAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT102 is further configured to receive data (e.g., voice input data) from the media playback system 100 and correspondingly transmit commands and / or media content to the media playback system 100.

[0045] The cloud network 102 comprises computing devices 106 (identified separately as a first computing device 106a, a second computing device 106b, and a third computing device 106c). The computing devices 106 can comprise individual computers or servers, such as, for example, a media streaming service server storing audio and / or other media content, a voice service server, a social media server, a media playback system control server, etc. In some embodiments, one or more of the computing devices 106 comprise modules of a single computer or server. In certain embodiments, one or more of the computing devices 106 comprise one or more modules, computers, and / or servers. Moreover, while the cloud network 102 is described above in the context of a single cloud network, in some embodiments the cloud network 102 comprises a plurality of cloud networks comprising communicatively coupled computing devices. Furthermore, while the cloud network 102 is shown in FIG. IB as having three of the computing devices 106, in some embodiments, the cloud network 102 comprises fewer (or more than) three computing devices 106.

[0046] The media playback system 100 is configured to receive media content from the networks 102 via the links 103. The received media content can comprise, for example, a Uniform Resource Identifier (URI) and / or a Uniform Resource Locator (URL). For instance, in some examples, the media playback system 100 can stream, download, or otherwise obtain data from a URI or a URL corresponding to the received media content. A network 104 communicatively couples the links 103 and at least a portion of the devices (e.g., one or more of the playback devices 110, NMDs 120, and / or control devices 130) of the media playback system 100. The network 104 can include, for example, a wireless network (e.g., a WIFI network, a BLUETOOTH network, a Z-WAVE network, a ZIGBEE network, and / or other suitable wireless communication protocol network) and / or a wired network (e.g., a network comprising Ethernet, Universal Serial Bus (USB), and / or another suitable wired communication). As those of ordinary skill in the art will appreciate, as used herein, “WIFI” can refer to several different communication protocols including, for example, Institute of Electrical and Electronics Engineers (IEEE) 802.11a, 802.11b, 802.11g, 802.1 In, 802.1 lac, 802. Had, 802.11af, 802. Hah, 802.1 lai, 802.1 laj, 802.11aq, 802.11ax, 802. Hay, 802.15, etc. transmitted at 2.4 Gigahertz (GHz), 5 GHz, 6 GHz, and / or another suitable frequency.

[0047] In some embodiments, the network 104 comprises a dedicated communication network that the media playback system 100 uses to transmit messages between individual devicesAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT and / or to transmit media content to and from media content sources (e.g., one or more of the computing devices 106). In certain embodiments, the network 104 is configured to be accessible only to devices in the media playback system 100, thereby reducing interference and competition with other household devices. In other embodiments, however, the network 104 comprises an existing household or commercial facility communication network (e.g., a household or commercial facility WIFI network). In some embodiments, the links 103 and the network 104 comprise one or more of the same networks. In some aspects, for example, the links 103 and the network 104 comprise a telecommunication network (e.g., an LTE network, a 5G network, etc.). Moreover, in some embodiments, the media playback system 100 is implemented without the network 104, and devices comprising the media playback system 100 can communicate with each other, for example, via one or more direct connections, PANs, telecommunication networks, and / or other suitable communication links. The network 104 may be referred to herein as a “local communication network” to differentiate the network 104 from the cloud network 102 that couples the media playback system 100 to remote devices, such as cloud servers that host cloud services.

[0048] In some embodiments, audio content sources may be regularly added or removed from the media playback system 100. In some embodiments, for example, the media playback system 100 performs an indexing of media items when one or more media content sources are updated, added to, and / or removed from the media playback system 100. The media playback system 100 can scan identifiable media items in some or all folders and / or directories accessible to the playback devices 110, and generate or update a media content database comprising metadata (e.g., title, artist, album, track length, etc.) and other associated information (e.g., URIs, URLs, etc.) for each identifiable media item found. In some embodiments, for example, the media content database is stored on one or more of the playback devices 110, network microphone devices 120, and / or control devices 130.

[0049] In the illustrated embodiment of FIG. IB, the playback devices 1101 and 110m comprise a group 107a. The playback devices 1101 and 110m can be positioned in different rooms and be grouped together in the group 107a on a temporary or permanent basis based on user input received at the control device 130a and / or another control device 130 in the media playback system 100. When arranged in the group 107a, the playback devices 1101 and 110m can be configured to play back the same or similar audio content in synchrony from one or more audio content sources. In certain embodiments, for example, the group 107a comprises a bonded zone in which the playback devices 1101 and 110m comprise left audio and right audio channels, respectively, of multi-channel audio content, thereby producing or enhancing a stereo effect ofAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT the audio content. In some embodiments, the group 107a includes additional playback devices110. In other embodiments, however, the media playback system 100 omits the group 107a and / or other grouped arrangements of the playback devices 110. Additional details regarding groups and other arrangements of playback devices are described in further detail below with respect to FIGS. II through IM.

[0050] The media playback system 100 includes the NMDs 120a and 120b, each comprising one or more microphones configured to receive voice utterances from a user. In the illustrated embodiment of FIG. IB, the NMD 120a is a standalone device and the NMD 120b is integrated into the playback device HOn. The NMD 120a, for example, is configured to receive voice input 121 from a user 123. In some embodiments, the NMD 120a transmits data associated with the received voice input 121 to a voice assistant service (VAS) configured to (i) process the received voice input data and (ii) facilitate one or more operations on behalf of the media playback system 100.

[0051] In some aspects, for example, the computing device 106c comprises one or more modules and / or servers of a VAS (e.g., a VAS operated by one or more of SONOS, AMAZON, GOOGLE, APPLE, MICROSOFT, etc.). The computing device 106c can receive the voice input data from the NMD 120a via the network 104 and the links 103.

[0052] In response to receiving the voice input data, the computing device 106c processes the voice input data (i.e., “Play Hey Jude by The Beatles”), and determines that the processed voice input includes a command to play a song (e.g., “Hey Jude”). In some embodiments, after processing the voice input, the computing device 106c accordingly transmits commands to the media playback system 100 to play back “Hey Jude” by the Beatles from a suitable media service (e.g., via one or more of the computing devices 106) on one or more of the playback devices 110. In other embodiments, the computing device 106c may be configured to interface with media services on behalf of the media playback system 100. In such embodiments, after processing the voice input, instead of the computing device 106c transmitting commands to the media playback system 100 causing the media playback system 100 to retrieve the requested media from a suitable media service, the computing device 106c itself causes a suitable media service to provide the requested media to the media playback system 100 in accordance with the user’s voice utterance. b. Suitable Playback Devices

[0053] FIG. 1C is a block diagram of the playback device 110a comprising an input / output111. The input / output 111 can include an analog I / O I l la (e.g., one or more wires, cables, and / or other suitable communication links configured to carry analog signals) and / or a digitalAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCTI / O 11 lb (e.g., one or more wires, cables, or other suitable communication links configured to carry digital signals). In some embodiments, the analog I / O I l la is an audio line-in input connection comprising, for example, an auto-detecting 3.5mm audio line-in connection. In some embodiments, the digital I / O 111b comprises a Sony / Philips Digital Interface Format (S / PDIF) communication interface and / or cable and / or a Toshiba Link (TOSLINK) cable. In some embodiments, the digital I / O 111b comprises a High-Definition Multimedia Interface (HDMI) interface and / or cable. In some embodiments, the digital I / O 111b includes one or more wireless communication links comprising, for example, a radio frequency (RF), infrared, WIFI, BLUETOOTH, or another suitable communication link. In certain embodiments, the analog EO I l la and the digital EO 111b comprise interfaces (e.g., ports, plugs, jacks, etc.) configured to receive connectors of cables transmitting analog and digital signals, respectively, without necessarily including cables.

[0054] The playback device 110a, for example, can receive media content (e.g., audio content comprising music and / or other sounds) from a local audio source 105 via the input / output 111 (e.g., a cable, a wire, a PAN, a BLUETOOTH connection, an ad hoc wired or wireless communication network, and / or another suitable communication link). The local audio source 105 can comprise, for example, a mobile device (e.g., a smartphone, a tablet, a laptop computer, etc.) or another suitable audio component (e.g., a television, a desktop computer, an amplifier, a phonograph (such as n LP turntable), a Blu-ray player, a memory storing digital media files, etc.). In some aspects, the local audio source 105 includes local music libraries on a smartphone, a computer, a networked-attached storage (NAS), and / or another suitable device configured to store media files. In certain embodiments, one or more of the playback devices 110, NMDs 120, and / or control devices 130 comprise the local audio source 105. In other embodiments, however, the media playback system omits the local audio source 105 altogether. In some embodiments, the playback device 110a does not include an input / output 111 and receives all audio content via the network 104.

[0055] The playback device 110a further comprises electronics 112, a user interface 113 (e.g., one or more buttons, knobs, dials, touch-sensitive surfaces, displays, touchscreens, etc.), and one or more transducers 114 (referred to hereinafter as “the transducers 114”). The electronics 112 are configured to receive audio from an audio source (e.g., the local audio source 105) via the input / output 111 or one or more of the computing devices 106a-c via the network 104 (FIG. IB), amplify the received audio, and output the amplified audio for playback via one or more of the transducers 114. In some embodiments, the playback device 110a optionally includes one or more microphones 115 (e.g., a single microphone, a plurality of microphones, aAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT microphone array) (hereinafter referred to as “the microphones 115”). In certain embodiments, for example, the playback device 110a having one or more of the optional microphones 115 can operate as an NMD configured to receive voice input from a user and correspondingly perform one or more operations based on the received voice input.

[0056] In the illustrated embodiment of FIG. 1C, the electronics 112 comprise one or more processors 112a (referred to hereinafter as “the processors 112a”), memory 112b, software components 112c, a network interface 112d, one or more audio processing components 112g (referred to hereinafter as “the audio components H2g”), one or more audio amplifiers 112h (referred to hereinafter as “the amplifiers 112h”), and power 112i (e.g., one or more power supplies, power cables, power receptacles, batteries, induction coils, Power-over Ethernet (POE) interfaces, and / or other suitable sources of electric power). In some embodiments, the electronics 112 optionally include one or more other components 112j (e.g., one or more sensors, video displays, touchscreens, battery charging bases, etc.).

[0057] The processors 112a can comprise clock-driven computing component(s) configured to process data, and the memory 112b can comprise a computer-readable medium (e.g., a tangible, non-transitory computer-readable medium loaded with one or more of the software components 112c) configured to store instructions for performing various operations and / or functions. The processors 112a are configured to execute the instructions stored on the memory 112b to perform one or more of the operations. The operations can include, for example, causing the playback device 110a to retrieve audio data from an audio source (e.g., one or more of the computing devices 106a-c (FIG. IB)), and / or another one of the playback devices 110. In some embodiments, the operations further include causing the playback device 110a to send audio data to another one of the playback devices 110a and / or another device (e.g., one of the NMDs 120). Certain embodiments include operations causing the playback device 110a to pair with another of the one or more playback devices 110 to enable a multi-channel audio environment (e.g., a stereo pair, a bonded zone, etc.).

[0058] The processors 112a can be further configured to perform operations causing the playback device 110a to synchronize playback of audio content with another of the one or more playback devices 110. As those of ordinary skill in the art will appreciate, during synchronous playback of audio content on a plurality of playback devices, a listener will preferably be unable to perceive time-delay differences between playback of the audio content by the playback device 110a and the other one or more other playback devices 110. Additional details regarding audio playback synchronization among playback devices can be found, for example, in U.S. Patent No. 8,234,395, which is incorporated by reference above.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0059] In some embodiments, the memory 112b is further configured to store data associated with the playback device 110a, such as one or more zones and / or zone groups of which the playback device 110a is a member, audio sources accessible to the playback device 110a, and / or a playback queue that the playback device 110a (and / or another of the one or more playback devices) can be associated with. The stored data can comprise one or more state variables that are periodically updated and used to describe a state of the playback device 110a. The memory 112b can also include data associated with a state of one or more of the other devices (e.g., the playback devices 110, NMDs 120, control devices 130) of the media playback system 100. In some aspects, for example, the state data is shared during predetermined intervals of time (e.g., every 5 seconds, every 10 seconds, every 60 seconds, etc.) among at least a portion of the devices of the media playback system 100, so that one or more of the devices have the most recent data associated with the media playback system 100.

[0060] The network interface 112d is configured to facilitate a transmission of data between the playback device 110a and one or more other devices on a data network such as, for example, the links 103 and / or the network 104 (FIG. IB). The network interface 112d is configured to transmit and receive data corresponding to media content (e.g., audio content, video content, text, photographs) and other signals (e.g., non-transitory signals) comprising digital packet data including an Internet Protocol (IP)-based source address and / or an IP -based destination address. The network interface 112d can parse the digital packet data such that the electronics 112 properly receive and process the data destined for the playback device 110a.

[0061] In the illustrated embodiment of FIG. 1C, the network interface 112d comprises one or more wireless interfaces 112e (referred to hereinafter as “the wireless interface 112e”). The wireless interface 112e (e.g., a suitable interface comprising one or more antennae) can be configured to wirelessly communicate with one or more other devices (e.g., one or more of the other playback devices 110, NMDs 120, and / or control devices 130) that are communicatively coupled to the network 104 (FIG. IB) in accordance with a suitable wireless communication protocol (e.g., WIFI, BLUETOOTH, LTE, etc.). In some embodiments, the network interface 112d optionally includes a wired interface 112f (e.g., an interface or receptacle configured to receive a network cable such as an Ethernet, a USB-A, USB-C, and / or Thunderbolt cable) configured to communicate over a wired connection with other devices in accordance with a suitable wired communication protocol. In certain embodiments, the network interface 112d includes the wired interface 112f and excludes the wireless interface 112e. In some embodiments, the electronics 112 exclude the network interface 112d altogether and transmitAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT and receive media content and / or other data via another communication path (e.g., the input / output 111).

[0062] The audio components 112g are configured to process and / or filter data comprising media content received by the electronics 112 (e.g., via the input / output 111 and / or the network interface 112d) to produce output audio signals. In some embodiments, the audio processing components 112g comprise, for example, one or more digital-to-analog converters (DACs), audio preprocessing components, audio enhancement components, digital signal processors (DSPs), and / or other suitable audio processing components, modules, circuits, etc. In certain embodiments, one or more of the audio processing components 112g can comprise one or more subcomponents of the processors 112a. In some embodiments, the electronics 112 omit the audio processing components 112g. In some aspects, for example, the processors 112a execute instructions stored on the memory 112b to perform audio processing operations to produce the output audio signals.

[0063] The amplifiers 112h are configured to receive and amplify the audio output signals produced by the audio processing components 112g and / or the processors 112a. The amplifiers 112h can comprise electronic devices and / or components configured to amplify audio signals to levels sufficient for driving one or more of the transducers 114. In some embodiments, for example, the amplifiers 112h include one or more switching or class-D power amplifiers. In other embodiments, however, the amplifiers 112h include one or more other types of power amplifiers (e.g., linear gain power amplifiers, class-A amplifiers, class-B amplifiers, class-AB amplifiers, class-C amplifiers, class-D amplifiers, class-E amplifiers, class-F amplifiers, class- G amplifiers, class H amplifiers, and / or another suitable type of power amplifier). In certain embodiments, the amplifiers 112h comprise a suitable combination of two or more of the foregoing types of power amplifiers. Moreover, in some embodiments, individual ones of the amplifiers 112h correspond to individual ones of the transducers 114. In other embodiments, however, the electronics 112 include a single one of the amplifiers 112h configured to output amplified audio signals to a plurality of the transducers 114. In some other embodiments, the electronics 112 omit the amplifiers 112h.

[0064] The transducers 114 (e.g., one or more speakers and / or speaker drivers) receive the amplified audio signals from the amplifier 112h and render or output the amplified audio signals as sound (e.g., audible sound waves having a frequency between about 20 Hertz (Hz) and 20 kilohertz (kHz)). In some embodiments, the transducers 114 can comprise a single transducer. In other embodiments, however, the transducers 114 comprise a plurality of audio transducers. In some embodiments, the transducers 114 comprise more than one type ofAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT transducer. For example, the transducers 114 can include one or more low frequency transducers (e.g., subwoofers, woofers), mid-range frequency transducers (e.g., mid-range transducers, mid-woofers), and one or more high frequency transducers (e.g., one or more tweeters). As used herein, “low frequency” can generally refer to audible frequencies below about 500 Hz, “mid-range frequency” can generally refer to audible frequencies between about 500 Hz and about 2 kHz, and “high frequency” can generally refer to audible frequencies above 2 kHz. In certain embodiments, however, one or more of the transducers 114 comprise transducers that do not adhere to the foregoing frequency ranges. For example, one of the transducers 114 may comprise a mid-woofer transducer configured to output sound at frequencies between about 200 Hz and about 5 kHz.

[0065] By way of illustration, Sonos, Inc. presently offers (or has offered) for sale certain playback devices including, for example, a “SONOS ONE,” “PLAY:1,” “PLAY:3,” “PLAYA,” “PLAYBAR,” “PLAYBASE,” “CONNECT: AMP,” “CONNECT,” “AMP,” “ARC,” “BEAM,” “PORT,” and “SUB .” Other suitable playback devices may additionally or alternatively be used to implement the playback devices of example embodiments disclosed herein. Additionally, one of ordinary skill in the art will appreciate that a playback device is not limited to the examples described herein or to Sonos product offerings. In some embodiments, for example, one or more playback devices 110 comprise wired or wireless headphones (e.g., over-the-ear headphones, on-ear headphones, in-ear earphones, etc.). In other embodiments, one or more of the playback devices 110 comprise a docking station and / or an interface configured to interact with a docking station for personal mobile media playback devices. In certain embodiments, a playback device may be integral to another device or component such as a television, an LP turntable, a lighting fixture, or some other device for indoor or outdoor use. In some embodiments, a playback device omits a user interface and / or one or more transducers. For example, FIG. ID is a block diagram of a playback device 1 lOp comprising the input / output 111 and electronics 112 without the user interface 113 or transducers 114.

[0066] FIG. IE is a block diagram of a bonded playback device 1 lOq comprising the playback device 110a (FIG. 1C) sonically bonded with the playback device HOi (e.g., a subwoofer) (FIG. 1 A). In the illustrated embodiment, the playback devices 110a and 1 lOi are separate ones of the playback devices 110 housed in separate enclosures. In some embodiments, however, the bonded playback device HOq comprises a single enclosure housing both the playback devices 110a and HOi. The bonded playback device HOq can be configured to process and reproduce sound differently than an unbonded playback device (e.g., the playback device 110aAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT of FIG. 1C) and / or paired or bonded playback devices (e.g., the playback devices 1101 and 110m of FIG. IB). In some embodiments, for example, the playback device 110a is a full-range playback device configured to render low frequency, mid-range frequency, and high frequency audio content, and the playback device 1 lOi is a subwoofer configured to render low frequency audio content. In some aspects, the playback device 110a, when bonded with the first playback device, is configured to render only the mid-range and high frequency components of a particular audio content, while the playback device 1 lOi renders the low frequency component of the particular audio content. In some embodiments, the bonded playback device HOq includes additional playback devices and / or another bonded playback device. c. Suitable Network Microphone Devices (NMDs)

[0067] FIG. IF is a block diagram of the NMD 120a (FIGS. 1 A and IB). The NMD 120a includes one or more voice processing components 124 (hereinafter “the voice components 124”) and several components described with respect to the playback device 110a (FIG. 1C) including the processors 112a, the memory 112b, and the microphones 115. The NMD 120a optionally comprises other components also included in the playback device 110a (FIG. 1C), such as the user interface 113 and / or the transducers 114. In some embodiments, the NMD 120a is configured as a media playback device (e.g., one or more of the playback devices 110), and further includes, for example, one or more of the audio components 112g (FIG. 1C), the amplifiers 112h, and / or other playback device components. In certain embodiments, the NMD 120a comprises an Internet of Things (loT) device such as, for example, a thermostat, alarm panel, fire and / or smoke detector, etc. In some embodiments, the NMD 120a comprises the microphones 115, the voice processing components 124, and only a portion of the components of the electronics 112 described above with respect to FIG. 1C. In some aspects, for example, the NMD 120a includes the processor 112a and the memory 112b (FIG. 1C), while omitting one or more other components of the electronics 112. In some embodiments, the NMD 120a includes additional components (e.g., one or more sensors, cameras, thermometers, barometers, hygrometers, etc.).

[0068] In some embodiments, an NMD can be integrated into a playback device. FIG. 1G is a block diagram of a playback device 1 lOr comprising an NMD 120d. The playback device 1 lOr can comprise many or all of the components of the playback device 110a and further include the microphones 115 and voice processing components 124 (FIG. IF). The playback device HOr optionally includes an integrated control device 130c. The control device 130c can comprise, for example, a user interface (e.g., the user interface 113 of FIG. 1C) configured to receive user input (e.g., touch input, voice input, etc.) without a separate control device. InAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT other embodiments, however, the playback device 11 Or receives commands from another control device (e.g., the control device 130a of FIG. IB).

[0069] Referring again to FIG. IF, the microphones 115 are configured to acquire, capture, and / or receive sound from an environment (e.g., the environment 101 of FIG. 1A) and / or a room in which the NMD 120a is positioned. The received sound can include, for example, vocal utterances, audio played back by the NMD 120a and / or another playback device, background voices, ambient sounds, etc. The microphones 115 convert the received sound into electrical signals to produce microphone data. The voice processing components 124 receive and analyze the microphone data to determine whether a voice input is present in the microphone data. The voice input can comprise, for example, an activation word followed by an utterance including a user request. As those of ordinary skill in the art will appreciate, an activation word is a word or other audio cue signifying a user voice input. For instance, in querying the AMAZON VAS, a user might speak the activation word "Alexa." Other examples include "Ok, Google" for invoking the GOOGLE VAS and "Hey, Siri" for invoking the APPLE VAS.

[0070] After detecting the activation word, voice processing components 124 monitor the microphone data for an accompanying user request in the voice input. The user request may include, for example, a command to control a third-party device, such as a thermostat (e.g., NEST thermostat), an illumination device (e.g., a PHILIPS HUE lighting device), or a media playback device (e.g., a SONOS playback device). For example, a user might speak the activation word “Alexa” followed by the utterance “set the thermostat to 68 degrees” to set a temperature in a home (e.g., the environment 101 of FIG. 1 A). The user might speak the same activation word followed by the utterance “turn on the living room” to turn on illumination devices in a living room area of the home. The user may similarly speak an activation word followed by a request to play a particular song, an album, or a playlist of music on a playback device in the home. d. Suitable Control Devices

[0071] FIG. 1H is a partial schematic diagram of the control device 130a (FIGS. 1A and IB). As used herein, the term “control device” can be used interchangeably with “controller” or “control system.” Among other aspects, the control device 130a is configured to receive user input related to the media playback system 100 and, in response, cause one or more devices in the media playback system 100 to perform an action(s) or operation(s) corresponding to the user input. In the illustrated embodiment, the control device 130a comprises a smartphone (e.g., an iPhone™ an Android phone, etc.) on which media playback system controller applicationAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT software is installed. In some embodiments, the control device 130a comprises, for example, a tablet (e.g., an iPad™), a computer (e.g., a laptop computer, a desktop computer, etc.), and / or another suitable device (e.g., a television, an automobile audio head unit, an loT device, etc.). In certain embodiments, the control device 130a comprises a dedicated controller for the media playback system 100. In other embodiments, as described above with respect to FIG. 1G, the control device 130a is integrated into another device in the media playback system 100 (e.g., one more of the playback devices 110, NMDs 120, and / or other suitable devices configured to communicate over a network).

[0072] The control device 130a includes electronics 132, a user interface 133, one or more speakers 134, and one or more microphones 135. The electronics 132 comprise one or more processors 132a (referred to hereinafter as “the processors 132a”), a memory 132b, software components 132c, and a network interface 132d. The processor 132a can be configured to perform functions relevant to facilitating user access, control, and configuration of the media playback system 100. The memory 132b can comprise data storage that can be loaded with one or more of the software components executable by the processor 132a to perform those functions. The software components 132c can comprise applications and / or other executable software configured to facilitate control of the media playback system 100. The memory 132b can be configured to store, for example, the software components 132c, media playback system controller application software, and / or other data associated with the media playback system 100 and the user.

[0073] The network interface 132d is configured to facilitate network communications between the control device 130a and one or more other devices in the media playback system 100, and / or one or more remote devices. In some embodiments, the network interface 132d is configured to operate according to one or more suitable communication industry standards (e.g., infrared, radio, wired standards including IEEE 802.3, wireless standards including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G, LTE, etc.). The network interface 132d can be configured, for example, to transmit data to and / or receive data from the playback devices 110, the NMDs 120, other ones of the control devices 130, one of the computing devices 106 of FIG. IB, devices comprising one or more other media playback systems, etc. The transmitted and / or received data can include, for example, playback device control commands, state variables, playback zone and / or zone group configurations. For instance, based on user input received at the user interface 133, the network interface 132d can transmit a playback device control command (e.g., volume control, audio playback control, audio content selection, etc.) from the control device 130a to one or more of the playback devicesAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT110. The network interface 132d can also transmit and / or receive configuration changes such as, for example, adding / removing one or more playback devices 110 to / from a zone, adding / removing one or more zones to / from a zone group, forming a bonded or consolidated player, separating one or more playback devices from a bonded or consolidated player, among others. Additional description of zones and groups can be found below with respect to FIGS. II through IM.

[0074] The user interface 133 is configured to receive user input and can facilitate control of the media playback system 100. The user interface 133 includes media content art 133a (e.g., album art, lyrics, videos, etc.), a playback status indicator 133b (e.g., an elapsed and / or remaining time indicator), media content information region 133c, a playback control region 133d, and a zone indicator 133e. The media content information region 133c can include a display of relevant information (e.g., title, artist, album, genre, release year, etc.) about media content currently playing and / or media content in a queue or playlist. The playback control region 133d can include selectable (e.g., via touch input and / or via a cursor or another suitable selector) icons to cause one or more playback devices in a selected playback zone or zone group to perform playback actions such as, for example, play or pause, fast forward, rewind, skip to next, skip to previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit cross fade mode, etc. The playback control region 133d may also include selectable icons to modify equalization settings, playback volume, and / or other suitable playback actions. In the illustrated embodiment, the user interface 133 comprises a display presented on a touch screen interface of a smartphone (e.g., an iPhone™ an Android phone, etc.). In some embodiments, however, user interfaces of varying formats, styles, and interactive sequences may alternatively be implemented on one or more network devices to provide comparable control access to a media playback system.

[0075] The one or more speakers 134 (e.g., one or more transducers) can be configured to output sound to the user of the control device 130a. In some embodiments, the one or more speakers comprise individual transducers configured to correspondingly output low frequencies, mid-range frequencies, and / or high frequencies. In some aspects, for example, the control device 130a is configured as a playback device (e.g., one of the playback devices 110). Similarly, in some embodiments the control device 130a is configured as an NMD (e.g., one of the NMDs 120), receiving voice commands and other sounds via the one or more microphones 135.

[0076] The one or more microphones 135 can comprise, for example, one or more condenser microphones, electret condenser microphones, dynamic microphones, and / or other suitableAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT types of microphones or transducers. In some embodiments, two or more of the microphones 135 are arranged to capture location information of an audio source (e.g., voice, audible sound, etc.) and / or configured to facilitate filtering of background noise. Moreover, in certain embodiments, the control device 130a is configured to operate as a playback device and an NMD. In other embodiments, however, the control device 130a omits the one or more speakers 134 and / or the one or more microphones 135. For instance, the control device 130a may comprise a device (e.g., a thermostat, an loT device, a network device, etc.) comprising a portion of the electronics 132 and the user interface 133 (e.g., a touch screen) without any speakers or microphones. e. Suitable Playback Device Configurations

[0077] FIGS. II through IM show example configurations of playback devices in zones and zone groups. Referring first to FIG. IM, in one example, a single playback device may belong to a zone. For example, the playback device 110g in the second bedroom 101c (FIG. 1 A) may belong to Zone C. In some implementations described below, multiple playback devices may be “bonded” to form a “bonded pair” which together form a single zone. For example, the playback device 1101 (e.g., a left playback device) can be bonded to the playback device 110m (e.g., a right playback device) to form Zone B. Bonded playback devices may have different playback responsibilities (e.g., channel responsibilities). In another implementation described below, multiple playback devices may be merged to form a single zone. For example, the playback device 1 lOh (e.g., a front playback device) may be merged with the playback device 1 lOi (e.g., a subwoofer), and the playback devices 1 lOj and 110k (e.g., left and right surround speakers, respectively) to form a single Zone D. In another example, the playback devices 110b and 1 lOd can be merged to form a merged group or a zone group 108b. The merged playback devices 110b and HOd may not be specifically assigned different playback responsibilities. That is, the merged playback devices 110b and 1 lOd may, aside from playing audio content in synchrony, each play audio content as they would if they were not merged.

[0078] Each zone in the media playback system 100 may be provided for control as a single user interface (UI) entity. For example, Zone A may be provided as a single entity named Master Bathroom. Zone B may be provided as a single entity named Master Bedroom. Zone C may be provided as a single entity named Second Bedroom.

[0079] Playback devices that are bonded may have different playback responsibilities, such as responsibilities for certain audio channels. For example, as shown in FIG. II, the playback devices 1101 and 110m may be bonded so as to produce or enhance a stereo effect of audio content. In this example, the playback device 1101 may be configured to play a left channelAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT audio component, while the playback device 110m may be configured to play a right channel audio component. In some implementations, such stereo bonding may be referred to as “pairing.”

[0080] Additionally, bonded playback devices may have additional and / or different respective speaker drivers. As shown in FIG. 1 J, the playback device 1 lOh named Front may be bonded with the playback device 1 lOi named SUB. The Front device 1 lOh can be configured to render a range of mid to high frequencies and the SUB device 1 lOi can be configured to render low frequencies. When unbonded, however, the Front device 11 Oh can be configured to render a full range of frequencies. As another example, FIG. IK shows the Front and SUB devices 1 lOh and 1 lOi further bonded with Left and Right playback devices 1 lOj and 110k, respectively. In some implementations, the Left and Right devices HOj and 110k can be configured to form surround or “satellite” channels of a home theater system. The bonded playback devices 1 lOh, 1 lOi, 1 lOj, and 110k may form a single Zone D (FIG. IM).

[0081] Playback devices that are merged may not have assigned playback responsibilities and may each render the full range of audio content the respective playback device is capable of. Nevertheless, merged devices may be represented as a single UI entity (i.e., a zone, as discussed above). For instance, the playback devices 110a and 11 On in the master bathroom have the single UI entity of Zone A. In one embodiment, the playback devices 110a and 1 lOn may each output the full range of audio content each respective playback devices 110a and 11 On are capable of, in synchrony.

[0082] In some embodiments, an NMD is bonded or merged with another device so as to form a zone. For example, the NMD 120b may be bonded with the playback device I lOe, which together form Zone F, named Living Room. In other embodiments, a stand-alone network microphone device may be in a zone by itself. In other embodiments, however, a stand-alone network microphone device may not be associated with a zone. Additional details regarding associating network microphone devices and playback devices as designated or default devices may be found, for example, in U.S. Patent No. 10,499,146 filed February 21, 2017 and titled “VOICE CONTROL OF A MEDIA PLAYBACK SYSTEM,” which is incorporated herein by reference in its entirety for all purposes.

[0083] Zones of individual, bonded, and / or merged devices may be grouped to form a zone group. For example, referring to FIG. IM, Zone A may be grouped with Zone B to form a zone group 108a that includes the two zones. Similarly, Zone G may be grouped with Zone H to form the zone group 108b. As another example, Zone A may be grouped with one or more other Zones C-I. The Zones A-I may be grouped and ungrouped in numerous ways. ForAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT example, three, four, five, or more (e.g., all) of the Zones A-I may be grouped. When grouped, the zones of individual and / or bonded playback devices may play back audio in synchrony with one another, as described in previously referenced U.S. Patent No. 8,234,395. Playback devices may be dynamically grouped and ungrouped to form new or different groups that synchronously play back audio content.

[0084] In various implementations, the zones in an environment may be the default name of a zone within the group or a combination of the names of the zones within a zone group. For example, Zone Group 108b can be assigned a name such as “Dining + Kitchen”, as shown in FIG. IM. In some embodiments, a zone group may be given a unique name selected by a user.

[0085] Certain data may be stored in a memory of a playback device (e.g., the memory 112b of FIG. 1C) as one or more state variables that are periodically updated and used to describe the state of a playback zone, the playback device(s), and / or a zone group associated therewith. The memory may also include the data associated with the state of the other devices of the media system, and shared from time to time among the devices so that one or more of the devices have the most recent data associated with the system.

[0086] In some embodiments, the memory may store instances of various variable types associated with the states. Variable instances may be stored with identifiers (e.g., tags) corresponding to type. For example, certain identifiers may be a first type “al” to identify playback device(s) of a zone, a second type “bl” to identify playback device(s) that may be bonded in the zone, and a third type “cl” to identify a zone group to which the zone may belong. As a related example, identifiers associated with the second bedroom 101c may indicate that the playback device is the only playback device of the Zone C and not in a zone group. Identifiers associated with the Den may indicate that the Den is not grouped with other zones but includes bonded playback devices 11 Oh- 110k. Identifiers associated with the Dining Room may indicate that the Dining Room is part of the Dining + Kitchen zone group 108b and that devices 110b and 1 lOd are grouped (FIG. IL). Identifiers associated with the Kitchen may indicate the same or similar information by virtue of the Kitchen being part of the Dining + Kitchen zone group 108b. Other example zone variables and identifiers are described below.

[0087] In yet another example, the memory may store variables or identifiers representing other associations of zones and zone groups, such as identifiers associated with Areas, as shown in FIG. IM. An area may involve a cluster of zone groups and / or zones not within a zone group. For instance, FIG. IM shows an Upper Area 109a including Zones A-D and I, and a Lower Area 109b including Zones E-I. In one aspect, an Area may be used to invoke a cluster of zone groups and / or zones that share one or more zones and / or zone groups of another cluster. InAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT another aspect, this differs from a zone group, which does not share a zone with another zone group. Further examples of techniques for implementing Areas may be found, for example, in U.S. Patent No. 10,712,997 filed August 21, 2017, and titled “Room Association Based on Name,” and U.S. Patent No. 8,483,853 filed September 11, 2007, and titled “Controlling and manipulating groupings in a multi-zone media system.” Each of these patents is incorporated herein by reference in its entirety. In some embodiments, the media playback system 100 may not implement Areas, in which case the system may not store variables associated with Areas. III. Examples of Speech Enhancement Techniques

[0088] As discussed above, there are numerous instances in which in case be desirable to enhance the clarity and / or intelligibility of dialogue (speech audio) in a mixed audio stream (e.g., a soundtrack that includes overlapping speech audio and non-speech audio). Some approaches to speech enhancement involve separating the speech audio from non-speech audio such that different audio processing can be applied to the two groups. For example, the volume of the speech audio can be increased while the volume of the non-speech audio can be decreased. Thus, when the two are re-combined, the speech audio is louder and therefore may be easier for a listener to hear over the reduced background sound.

[0089] Referring to FIG. 2, there is illustrated an example of this approach. FIG. 2 illustrates an audio processing chain 200 that may be implemented by a playback device (e.g., any of the playback devices 110 or NMDs 120 described above). An input signal 202 is input to the audio processing chain 200. The input signal 202 may represent a mixed audio soundtrack, such as a soundtrack associated with displayed video (e.g., a movie or television series) or an audio narration that includes both speech and non-speech background sound. Speech extraction techniques may operate in the frequency domain. Accordingly, in some examples, at operation 204, a short-time Fourier transform (STFT) is applied to the input signal 202 to produce a corresponding input frequency spectrum 206. In the digital domain, the input frequency spectrum 206 may be represented as a two-dimensional matrix of signal magnitude and frequency. For example, based on a selected resolution of the STFT, the input signal 202 is divided into a plurality of frequency bins (e.g., 256, 512, etc.) and the signal magnitude in each frequency bin is recorded as a digital value. In some examples, the STFT at operation 204, and other processing operations of the audio processing chain 200, may be implemented by the electronics 112 (e.g., the processor(s) 112a, audio processing components 112g, and / or software components 112c) of a playback device 110. In some examples, a frequency sample size of the STFT can be adjusted based on one of more attributes of the audio signal. ForAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT example, certain audio signals may have more closely spaced frequency content and may benefit from a higher resolution (e.g., more frequency bins) transform.

[0090] To extract the speech audio from the input signal 202, an adaptive controller 208 processes the input frequency spectrum 206 to produce another matrix referred to as a speech mask 210. For example, the speech mask 210 may be represented by a two-dimensional matrix having the same number of frequency bins as the input frequency spectrum 206, with a logical 1 or 0 value associated with each frequency bin indicating speech (e.g., 1) or no speech (e.g., 0). Inverting the speech mask210 produces a noise mask 212. The adaptive controller 208 may be implemented by the electronics 112, for example, by the processor(s) 112a and / or the software components 112c. In some examples, the adaptive controller 208 includes a parametric machine-learning model, as described further below.

[0091] Multiplying the input frequency spectrum 206 by the speech mask 210 produces a speech signal 214. In examples, the speech mask 210 “masks out” the non-speech content of the original input frequency spectrum 206, such that the speech signal 214 contains only, or predominantly, the speech content / information from the original input signal 202. Similarly, multiplying the input frequency spectrum 206 by the noise mask 212 produces a background signal 216 from which the speech content has largely been eliminated. Accordingly, the speech signal 214 and the background signal 216 can be separately processed, as described further below.

[0092] A remixing operation 218 can then combine the, optionally processed, speech signal 214 and background signal 216 to produce a new signal 220 for which the speech is enhanced with respect to the original signal 202. For example, the magnitude of the speech content can be relatively increased while the magnitude of the background sound can be relatively decreased (with respect to the original signal 202 and to each other). The new signal 220 can be converted back into the time domain by applying an inverse short-time Fourier transform (iSTFT) at operation 222.

[0093] According to certain approaches, extraction of speech audio from the input frequency spectrum 206 to produce the speech mask 210 can be performed based on estimated frequency content of speech. For example, it is known that human speech typically occurs in certain frequency bands. Accordingly, a system can be configured to produce the speech mask 210 based on this information (e.g., frequency bins corresponding to those frequencies generally known to be associated with human speech can be given a logical 1 and all other frequency bins given a logical 0). Thus, this approach does not rely on detecting speech content, but rather uses generalized frequency content information. However, as a result, this approach may lackAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT accuracy (e.g., not all speech contains the complete range of speech-associated frequency content) and also operates all the time, whether speech is in fact present in a particular frame of the input signal 202 or not. This “always on” approach can negatively impact the quality of the audio when no dialogue is present in an audio frame. For example, if the input signal 202 contains well-mixed, balanced audio (with no dialogue), applying speech enhancement audio processing can disrupt the balance and degrade the sound experience. For example, audio content at those frequencies corresponding to typical speech-associated frequencies may be amplified relative to the audio content at other frequencies, which can cause the soundtrack to sound distorted or otherwise alter the intended sound experience.

[0094] Furthermore, in some instances, speech overlaps in frequency with certain background sounds, and can therefore be at least partially masked or covered by the background sounds. For example, in a video scene where two characters are walking through long grass, high frequency detail of sound representing the moving grass may occupy a similar range to the sibilance of the dialogue between the two characters. As a result, an approach to speech enhancement that is based solely on estimated frequencies associated with speech may boost the grass sounds in addition to the dialogue, and therefore fail to improve the intelligibility of the dialogue.

[0095] Accordingly, as described above, certain examples apply speech detection techniques to separate dialogue from other sounds in a soundtrack. Using such techniques, the speech mask 210 can be produced based on a more accurate representation of the speech content of a particular frame of the input signal 202. According to certain examples, the adaptive controller 208 implements machine learning approaches to identify speech content in the input frequency spectrum 206 and produce a corresponding speech mask 210. In some examples, this process can be applied on a frame by frame basis to the input signal 202, such that speech enhancement techniques (described further below) can be applied adaptively (e.g., based on an estimation of the actual speech content in any given frame) and in real-time. When speech is detected, speech enhancement processing can be applied to “uncover” dialogue from background sounds and improve the clarity of dialogue. However, when no speech is detected, the original audio mix of the input signal 202 remains unchanged. Thus, according to certain examples, the clarity of dialogue can be improved without negatively impacting the overall sound experience, particularly when no dialogue is present in the soundtrack.

[0096] According to certain examples, the adaptive controller 208 can be configured to implement a model predictive controller that runs one or more parameterized machine learning models that can be trained to identify speech in the input signal 202. In some examples, theAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT model(s) operate on data representing the input frequency spectrum 206 corresponding to the input signal 202; however, in other examples, the STFT operation 204 can be omitted, and the model(s) can operate based on data acquired from, or directly representing, the time-domain input signal 202.

[0097] FIG. 3 illustrates a machine learning system 300 that can be used to implement speech extraction techniques according to certain examples. The machine learning system 300 may be an implementation of, or part of, the adaptive controller 208. In the example of FIG. 3, a model predictive controller 302 operates on input data 304 and its operation is controlled by an optimizer 306. The model predictive controller (MPC) 802 includes a model 308, which may be a parameterized (or parametric) machine learning model, as discussed above. In some examples, the MPC 302 further includes a data sampler 310, user preferences 312, and a confidence element 314. However, in other examples, the MPC 302 may omit the user preferences 312 and / or the confidence element 314. In some examples, the confidence element 314 may apply two threshold values, namely an uncertainty threshold 316 and a decision threshold 318, each of which is discussed further below. The confidence element 314 allows the system 300 to accommodate uncertainty in the prediction (e.g., by using confidence indicators, as described below), which can lead to improved performance. The system 300 may be implemented, in whole or in part, on one or more playback devices 110 within the MPS 100. The system 300 may be implemented in software or using any combination of hardware and software capable of performing the functions disclosed herein.

[0098] In examples, the model predictive controller 302 runs the model 308 based on parameters associated with one or more features extracted from the input data 304 to detect speech in the input signal 202. In some examples, the input data 304 is digital data that represents at least a portion of the input signal 202 and / or the input frequency spectrum 206. For example, the input data 304 can include the input frequency spectrum 206, or a portion thereof, represented as a two-dimensional matrix, as described above. In some examples, the input data 304 can further include context data describing one or more aspects of the playback environment in which audio content from the input signal 202 is to be played back by one or more playback devices. For example, the context data can include information such as whether the playback device operating the machine learning system 300 is currently part of a bonded group, and / or a source and / or type of the input signal (e.g., whether the input signal 202 represents music content, a podcast or other narrative audio content, or home theater audio content that is associated with corresponding video). In some instances, the input data 304 mayAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT include labeled and / or unlabeled training data that can be used to train the model 308, as described further below.

[0099] In examples, the data sampler 310 intakes the input data 304 and extracts one or more input features to be used by the model 308, as described further below.

[0100] The model 308 uses the input data 304 to generate a set of parameters which yield a generalized function capable of predicting one or more particular output values (e.g., whether or not speech is likely present in a current frame of the input signal 202) based on new input data 304. Parameters are variables that “belong” to the model 308 in that the trained model is represented by the model parameters. In contrast, hyperparameters are higher-level variables that affect the learning process and, thus, the values of the model parameters of the trained model 308. In some examples, training the model 308 involves choosing hyperparameters that the learning process uses to generate parameters that correctly map the input features (independent variables) to the labels (dependent variables) such that the model 308 produces predictions (e.g., speech or no speech) with reasonable accuracy.

[0101] In the example illustrated in FIG. 3, the system 300 includes the optimizer 306 that operates based on one or more hyperparameters to optimize performance of the MPC 302. In some examples, the optimizer 306 selects hyperparameters for use during training of the model 308. Hyperparameters may include variables that determine characteristics such as an architecture of the model 308 (e.g., kernel selection or type of model (e.g., Gaussian mixture model, hidden Markov model, neural network, etc.)), how the model 308 is applied, and / or variables that affect an optimization process used by the optimizer 306. A hyperparameter can, for example, take the form of a single continuous scalar variable or a discrete categorical variable (e.g., which model to use). Selection of hyperparameters has a significant impact on the performance (e.g., accuracy of predictions) of the trained model 308. Accordingly, in some examples, the optimizer 306 applies an optimization process to select the best hyperparameters for training the model 308. In some examples, this optimization process involves testing the performance of the system 300 on a validation dataset and adjusting the hyperparameters to produce an optimal result. Thus, the optimizer 306 may select certain hyperparameters, train a first model 308, and test the first model 308 using the validation data. The optimizer 306 may then tune the hyperparameters, train a second model 308, test the second model 308 using the validation data, and compare the performance to determine which hyperparameters produced a better result in the trained models. This process can be repeated to find optimal hyperparameters. In some examples, the optimizer 306 may use a grid search involving a field of combinations of hyperparameter values. In other examples, the optimizer 306 may apply aAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT gradient descent optimization or a gradient-free optimization method, such as Bayesian optimization, or some combination thereof. As noted above, in some examples, the choice of optimization process can be a hyperparameter itself.

[0102] Thus, hyperparameters are “external” to the model 308 since they cannot be changed by the model during training, although they may be tuned by the optimizer 306 to control the training of the model 308. As described above, a hyperparameter selected by the optimizer 306 can include a set of model parameters, as well as values that define the model architecture itself. In contrast, the model parameters are internal to the model 308 and their values are learned or estimated based on the input data 304 during training as the model 308 tries to learn the mapping between the input features and the labels.

[0103] According to certain examples, the model 308 can be a neural network, for example, a deep neural network (DNN) model. An artificial neural network (ANN) is based on a collection of connected nodes (artificial “neurons”), with each connection allowing a signal to be transmitted from one neuron to another. A receiving neuron can process received signals and transmit the processed signals to other downstream neurons. Neurons may have state, generally represented by real numbers (e.g., between 0 and 1). Neurons and connections may also have a weight that varies as learning proceeds, which can increase or decrease the strength of the signal that the neuron / connection sends downstream. Generally, the neurons are organized in layers, and different layers may perform different kinds of transformations on their inputs. Deep learning models may be based on multi-layered neural networks in which a hierarchy of layers is used to transform input data into a more suitable representation for a classification process to operate on. For speech recognition, for example, the classification process may operate on the transformed data to extract features that allow the model 308, based on its training, to identify probabilities that the extracted features match features associated with speech patterns. The term “deep” in deep neural network refers to the number of layers through which the data is transformed. In some examples, a deep (as opposed to shallow) neural network has a credit assignment path (CAP) depth of at least 2. The CAP is the chain of transformations from input to output and describes potentially causal connections between the input and the output.

[0104] In some examples, a DNN model 308 can be implemented as a recurrent neural network (RNN). An RNN is a network (optionally bi-directional) in which the output from some nodes can affect subsequent input to the same nodes. Information in the model can loop or cycle through one or more layers multiple times during flow from input nodes (an input layer),Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT through any hidden nodes (one or more hidden layers), to output nodes (an output layer). An RNN may have an infinite impulse response or a finite impulse response.

[0105] An RNN can use internal states of the neurons / nodes (memory) to process arbitrary sequences of input data. In certain examples, an RNN model 308 can be trained by the optimizer 306, for example, by applying a gradient descent optimization combined with backpropagation through time to compute the gradients that are used to change the weights associated with nodes / connections in the model. However, a problem that can arise when training an RNN using back-propagation is that the long-term gradients that are back- propagated can “vanish,”, meaning they can tend to zero due to very small numbers creeping into the computations, causing the model to effectively stop learning. Accordingly, in some examples, the RNN can be configured to use long short-term memory (LSTM) units to at least partially solve the vanishing gradient problem, because LSTM units allow gradients to also flow with little to no attenuation. An LSTM unit can keep track of arbitrary long-term dependencies in the input sequences. The LSTM architecture includes a memory module in the ANN that learns when to remember certain information (e.g., identifies information that might be needed later on in a sequence) and when to forget the information (e.g., identifies when that information is no longer needed).

[0106] In some examples, an LSTM unit comprises a cell, an input gate, an output gate, and a forget gate. The cell is a memory cell that stores values over arbitrary time intervals. The three gates regulate the flow of information into and out of the cell. Forget gates control what information to discard from a previous state by assigning a previous state, compared to a current input, a value between 0 and 1. For example, a (rounded) value of 1 may cause the cell to retain the information, whereas a (rounded) value of 0 may cause the cell to discard the information. Input gates control which pieces of new information to store in the current state, for example, by using the same system as used by the forget gates. Output gates control which pieces of information in the current state to output by assigning a value from 0 to 1 to the information, considering the previous and current states. Selectively outputting certain information from the current state allows the LSTM network to maintain useful, long-term dependencies that the model incorporates to make predictions.

[0107] Thus, in some examples, the model 308 can be implemented using an RNN, LSTM- based RNN, or other type of DNN model. For example, the model 308 can be implemented using a convolutional neural network (CNN). A CNN is a feed-forward neural network having a finite impulse response. The CNN learns features via filter (or kernel) optimization. The hidden layer(s) of a CNN include one or more layers that perform convolutions. In someAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT examples, a convolution layer performs a dot product of the convolution kernel with the input matrix of the layer. The convolution operation generates a feature map, which in turn contributes to the input of the next layer. In some applications, transformers can be used to replace RNN’s however, for some use cases, the implementation of transformers may not be practical due to the computational requirements of transformers.

[0108] In other examples, the model 308 can be implemented using a machine learning approach other than a DNN. For example, the model 308 may be implemented using a Gaussian mixture model (GMM) kernel or a hidden Markov model (HMM) kernel. Further, in some examples, the MPC 302 may include a plurality of different models 308, including, for example, one or more DNN models and / or one or more other models, such as a GMM and / or HMM, for example. In such instances, the MPC 302 may select which one or more models 308 to process a given frame of audio data included in the input data 304. For example, as described above, the input data 304 can include context data, which may identify a particular type of audio content that is to be processed (e.g., home theater content, music content, etc.). Accordingly, in some examples, the MPC 302 can select which model 308 to use based on the type of audio content to be processed. In some examples, the user preferences 312 and / or the confidence element 314 can influence selection of a particular model 308, as described further below.

[0109] Still referring to FIG. 8, as discussed above, in some examples, the MPC 302 incorporates uncertainty through the use of the confidence element 314. For example, the output from the model 308 may be a probability (e.g., the probability that a particular frequency bin of the frequency spectrum 306 for a current frame of the input signal 202 contains dialogue) and therefore has a built-in measure of uncertainty, or “confidence metric.” In one example, the uncertainty threshold 318 dictates a value at which the uncertainty in the model output is sufficiently low for the MPC 302 to trust the model prediction. For example, if the model output (prediction) is a 60% probability that the current audio frame contains dialogue in one or more frequency bins, the uncertainty (40% in this case) may be too high for the MPC 302 to trust the model prediction. This may indicate a need for re-training of the model 308 or selection of a different type of model, for example. In some examples, the decision threshold 316 dictates a value at which the uncertainty in the model output is sufficiently low for the MPC 302 to take a certain action based on the model output. This action may include causing speech enhancement processing to be applied, as described further below.

[0110] In some instances, the uncertainty threshold 318 and / or the decision threshold 316 can be hyperparameters that are applied (and optionally tuned) by the optimizer 306. For example,Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT the uncertainty threshold 318 and / or the decision threshold 316 may together define a trust region in which it is likely that acting on the model prediction will not result in undesirable system behavior (e.g., applying speech enhancement when no dialogue is present in the audio frame) and a negative user experience. In some examples, the optimizer 306 can be constrained to optimize the model parameters within this trust region set by the uncertainty threshold 318 and the decision threshold 316. In other examples, the uncertainty threshold 318 and / or the decision threshold 316 may directly affect the decision behavior of the MPC 802. For example, as described above, in some instances, the MPC 302 can be configured to automatically take an action (such as causing speech enhancement processing to be applied) if the uncertainty in the model prediction is below the limit set by the decision threshold 316 (e.g., below 10%, 5%, or 2% uncertainty, etc.).

[0111] According to certain examples, the MPC 302 may further acquire and store the user preferences 312. The user preferences 312 may include user-provided information regarding when and / or whether a user wants aspects of speech enhancement processing described herein to be applied. The user preferences 312 may be acquired as part of the input data 304 in some examples or may be separately acquired and stored. In some examples, a user may enter the user preferences 312 via a user interface 320, such as the user interface 133 on a control device 130, for example. In some examples, the user preferences 312 can be used to control how the hyperparameters are selected. For example, the optimizer 306 can be configured to optimize the model parameters within constraints set by the user preferences 312. User preferences 312 may also directly influence the behavior of the MPC 802, such as by constraining speech enhancement to certain playback environments (e.g., only when the playback device operating the machine learning system 300 is in a home theater bonded group). In some examples, the user preferences 312 can be used to set either or both of the decision threshold 316 and / or the uncertainty threshold 318. In this manner, a user can be provided with a wide degree of control over the behavior of the system 300 such that the system 300 can be configured in accord with an individual user’s own preferences and comfort level with system autonomy.

[0112] Furthermore, passive user feedback can be used to gauge the accuracy of the model predictions and adjust the model 308 to improve performance. For example, if the system 300 applies speech enhancement processing and the user takes no action, the MPC 302 may interpret that the prediction was correct. On the other hand, if the system 300 applies speech enhancement processing and the user turns off the speech enhancement processing, the MPC 302 may interpret that the prediction was incorrect. Thus, this passive user feedback can be acquired nearly continuously without bothering the user, since the feedback is acquired throughAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT the user’s natural interactions with the system, rather than through specific training-related tasks. This passive user feedback can be used to label the corresponding features associated with the input data set that produced the prediction and produce labeled training data. By retraining the model 308 with this labeled training data, the model 308 may produce similar predictions with higher or lower confidence metrics. In some examples, where the labeled training data is based on positive user feedback, the re-trained model 308 may produce a corresponding prediction based on similar input data 304 with a higher confidence metric (e.g., higher probability). In other examples, where the labeled training data is based on negative user feedback, the model 308 may be less likely to produce a corresponding prediction based on similar input data 304, or if it does, may produce the corresponding prediction with a lower confidence metric (e.g., lower probability).

[0113] As described above, the MPC 302 may operate the model 308 (or any one or more of multiple models 308) to extract speech / dialogue from an audio stream by recognizing the presence of the dialogue and producing the speech mask 210 based thereon. Accordingly, the model(s) 308 can be trained with a variety of different training data to allow the model(s) 308 to identify dialogue under different conditions. For example, the model(s) can be trained on a variety of mixed audio content soundtracks, such as movie soundtracks, television programming (e.g., series, news casts, sports programming, etc.), which can contain both different types of dialogue (e.g., shouting, whispering, normal speech, different voices, different language accents, different languages, etc.) and different background conditions / sound. Diverse training can allow the model(s) 308 to accurately extract speech content from a wide range of different input audio signals 202.

[0114] In some examples, an audio stream or soundtrack can contain multi-channel audio content. For example, as described above, in a home theater playback environment, a group of playback devices (e.g., the playback devices 1 lOh, 1 lOi, 1 lOj, and / or 110k) can be responsible for playing back various audio channels (e.g., a center channel, left and right front channels, one or more surround channels, etc.) of a multi-channel audio input signal. In such examples, dialogue may be present in one or more of the audio channels (e.g., the center channel), but not in others. Accordingly, in some instances, the input data 304 can be limited to one or more audio channels of the multi-channel audio signal (e.g., the center channel only). In other examples, the MPC 302 can be configured to separate the audio channels and optionally apply a different model 308 to different channels, or selectively apply the model 308 to one or more audio channels, but not others. In some examples, channel separation may be performed by the data sampler 310. In some examples, combined data from one or more channels of the multiAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT channel audio signal can be used as the input data 304 to produce one or more speech masks 210 that are applied, respectively, to one or more individual channels of the multi-channel audio signal. For example, data from the center channel, left and right front channels, and some or all of the surround channels can be used to detect speech and correlated audio in some or all of the channels and used to produce a speech mask 210 for one or more channels (e.g., the center channel). This approach may help to remove or reduce multi-channel effects that have been applied to speech (e.g., reverberation). In addition, this approach may be beneficial where speech is present across multiple channels and there is overlapping time and / or frequency content in the center channel.

[0115] As described above, in addition to extracting speech content to produce the speech mask 210, the audio processing chain 200 may involve applying speech enhancement techniques by processing the speech content (e.g., speech signal 214) differently from the non-speech content (e.g., the background signal 216). In some examples, the MPC 302 is used to perform dialogue extraction, and other audio processing circuitry (e.g., the audio processing components 112g) applies speech enhancement processing. For example, the audio processing circuitry can be configured to increase the volume (e.g., by changing amplification provided by the audio amplifiers 112h) of the speech signal 214 relative to the background signal 216. The audio processing circuitry may alternatively or additionally be configured to alter equalization settings, filtering, dynamic range compression, and / or other audio processing param eters / settings applied to the speech signal 214 and / or the background signal 216 so as to enhance the clarity of the dialogue. For example, the audio processing circuitry can be configured to apply transient shaping and / or other processes that enhance certain speech features (e.g., consonants, fricatives, and / or plosives) to improve the intelligibility of the speech. In other examples, the MPC 302 can be configured to control application of speech enhancement processing based on the output from the model 308. For example, based on the model 308 detecting dialogue in a given frame of the input signal 202, the MPC 302 can control audio processing circuitry to apply certain audio processing settings to one or more audio channels for speech enhancement, as described further below. In still further examples, the MPC 302 may operate the model 308 to apply speech enhancement audio processing. For example, the model 308 can be configured to predict one or more audio processing settings (e.g., gain, equalization, dynamic range compression, filtering, transient shaping or speech feature enhancement processes, etc.) for one or more audio channels based on detecting dialogue in the current audio frame, optionally in combination with context data (e.g., information indicating the type of soundtrack etc., as described above). In some examples, theAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCTMPC 302 can operate a first model 308 for dialogue detection / extraction, and a second model 308 for applying speech enhancement audio processing. In such instances, the first and second models 308 can be of the same or different types.

[0116] Referring now to FIG. 4 A, there is illustrated an example of a playback device 400 that can be configured to operate the machine learning system 300 for speech enhancement, according to certain aspects. The playback device 400 may be any of the playback devices 110 described above, for example. The playback device 400 includes one or more audio transducers 402 (e.g., the transducer(s) 114 described above). The playback device 400 includes a communication interface 404 (e.g., the network interface 112d including one or more wireless interfaces(s) 112e and / or one or more wired interface(s) 112f). In some examples, the playback device 400 receives an input audio signal 406 (e.g., the input signal 202) via the communication interface 404. The communication interface 404 may further allow the playback device 400 to communicate with one or more other playback device(s) 408. For example, the playback device 400 may be a home theater primary device, such as a soundbar, for example, and may communicate one or more audio channels of a multi-channel input audio signal 406 to one or more satellite playback devices 408. In some examples, the one or more other playback devices 408 can include a portable playback device, such as headphones or another type of portable playback device.

[0117] The playback device 400 further includes a digital signal processor 410 that comprises a speech processing device 412. The digital signal processor 410 may be an implementation of, be part of, or include, the processor(s) 112a, for example. The speech processing device 412 may include, or may be part of, the machine learning system 300. The digital signal processor 410 processes the input audio signal 406 to provide a speech content signal 414 and another signal 416. In some examples, the speech content signal 414 is the speech signal 214 described above and the other signal 416 is, includes, or is part of, the background signal 216 described above. In some examples in which the input audio signal 406 is a multi-channel audio signal, the other signal 416 may include one or more audio channels not processed for speech detection and / or enhancement, as described further below.

[0118] The playback device 400 further includes audio processing circuitry 418 that is configured to re-mix the speech content signal 414 and the other signal 216, optionally apply audio processing settings (e.g., equalization, filtering, dynamic range compression, gain, transient shaping or other speech feature enhancement, etc.) and provide one or more output audio signals to audio drivers associated with the transducer(s) 402. In some examples, one orAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT more audio output signals can also be provided from the audio processing circuitry 418 to the communication interface 404 for transmission to one or more satellite playback devices 408.

[0119] FIGS. 4B-D illustrate some examples of portable playback devices 408a-c, respectively. FIG. 4B is a front isometric view of an example of a portable playback device 408a a configured in accordance with aspects of the disclosed technology. As shown in FIGS. 4B and 4C, the portable playback devices 408a, 408b may be implemented as headphones to facilitate more private playback as compared with the out loud playback associated with other the satellite playback device(s) 408, such as portable playback device 408c or stationary playback devices 408, for example. In the example shown in FIG. 4B, the portable playback device 408a includes a housing 420a to support a pair of transducers 422a on or around a user’ s head over the user’ s ears. The portable playback device 408a also includes a user interface 424a with a touch-sensitive region to facilitate playback controls such as transport and / or volume controls. In some implementations, the user interface 424a may include respective touch- sensitive regions on the exterior of each earcup.

[0120] FIG. 4C is a front isometric view of an example of the portable playback device 408b, implemented as earbud-type headphones. As shown, the portable playback device 408b includes a housing 420b to support a pair of transducers 422b within a user’s ears. The portable playback device 408b also includes a user interface 424b with a touch-sensitive region to facilitate playback controls such as transport, volume, and / or swap controls. The portable playback device 408b can be in the form of wired or wireless earbuds, for example.

[0121] FIG. 4D is a front isometric view of an example of a portable playback device 408c. In certain examples, as compared with the headphones of FIGS. 4B and 4C, the portable playback device 408c may include one or more larger transducers (not shown) to facilitate out loud audio content playback. A speaker grill 426 covers the transducers. Relative to certain ones of the playback device(s) 110 discussed above, the portable playback device 408c may include less powerful amplifier(s) and / or smaller transducer(s) to balance battery life, sound output capability, and form factor (i.e., size, shape, and weight) of the portable playback device 408c. The portable playback device 408c includes a user interface 424c with a touch-sensitive region to facilitate playback controls such as transport, volume, and / or swap controls.

[0122] Turning now to FIG. 5, there is illustrated a flow diagram of audio processing 500 that can be performed by the playback device 400, according to certain examples. At operation 502, the playback device 400 receives at least one frame of the input audio signal 406. As described above, in some examples, dialogue / speech extraction and enhancement processing can be performed in real time on a frame by frame basis. However, in other examples, speechAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT extraction and enhancement processing may be performed on “blocks” of an audio stream that may include two or more frames.

[0123] At operation 504, the playback device 400 performs a dialogue extraction process on the current portion of the input audio signal received at operation 502. As described above, in some examples, operation 504 includes running the model 308 under control of the MPC 302 to extract the speech content from the current portion of the input audio signal. As also described above, in certain examples, the model 308 used in operation 504 is a DNN model trained for speech extraction.

[0124] Based on operation 504, the current portion of the input audio signal is separated into speech audio 506 (e.g., signals 214 or 414) and non-speech audio 508 (e.g., signals 216 or 416). Speech enhancement processing in certain examples includes an operation 410 of adjusting one or more characteristics of the speech audio 506. In some examples, the speech enhancement processing may further include an operation 412 of adjusting one or more characteristics of the non-speech audio 508. For example, at operation 510, the volume of the speech audio 506 can be increased and at operation 512, the volume of the non-speech audio 508 can be decreased. However, in other examples, at operation 510 the volume of the speech audio 506 can be increased and operation 512 can be omitted (e.g., the non-speech audio remains at an original volume level). Alternatively, at operation 512, the volume of the non-speech audio 508 can be decreased and operation 510 can be omitted (e.g., the speech audio remains at the original volume level). In further examples, operations 510 and / or 512 can include adjusting characteristics of the speech audio 506 and / or the non-speech audio 508, respectively, other than volume. The adjustment(s) performed at operations 510 and / or 512 can be accomplished by controlling settings of one or more audio processing components 112g and / or audio amplifiers 112h that may be part of the audio processing circuitry 418, or by controlling connections of components (e.g., amplifiers, filters, etc.) in the audio processing chain so as to achieve desired levels of amplification, filtering, balance, compression, etc.

[0125] According to certain examples, based on speech being detected at operation 504, the digital signal processor 410 can be configured to apply the various settings and / or connections according to a pre-programmed set of rules. As described above, in some examples, this control is executed by the MPC 302; however, in other examples, one or more other processors can implement the control based on a signal from the MPC 302 indicating that speech has been detected and speech enhancement processing is to be applied. In some examples, the MPC 302 can further supply information indicating a type of speech enhancement processing to be applied, as described further below. In further examples, based on speech being detected atAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT operation 504, the MPC 302 can run another model 308 to determine particular settings, connections, etc., to be implemented in order to achieve speech enhancement. The MPC 302 may then either effect control of the audio processing circuitry 418 or provide information to allow one or more other processors / controllers to configure the audio processing circuitry 418 appropriately to perform the speech enhancement processing (e.g., operations 510 and / or 512).

[0126] After speech enhancement processing is applied, the speech audio 506 and non-speech audio 508 can be recombined / re-mixed to provide an output audio signal 514. The output audio signal may then be provided to the transducer(s) 402 and / or one or more satellite playback devices 408 for playback. It will be appreciated that in some instances, other audio processing unrelated to speech enhancement may be applied to the output audio signal 514 and / or to the input audio signal 406 to allow for appropriate playback by the playback device 400 and / or one or more satellite playback devices 408.

[0127] It will further be appreciated that if no speech is detected by the model 308 at operation 504, the entire current portion of the input audio signal 406 may be considered the non-speech audio 508, operations 510 and 512 may be omitted, and the current portion of the input audio signal 406 may become the audio output signal 514. Thus, as described above, speech enhancement processing may be performed selectively and dynamically only when speech is detected, thereby avoiding the potential degradation of the sound experience that may occur if speech enhancement processing is applied to an audio signal that does not contain speech.

[0128] In certain examples, rather than re-mixing the speech audio 506 and the non-speech audio 508, the two signals can be directed to different transducers and / or playback devices for playback. For example, referring again to FIG. 4 A, in some examples, the speech content signal 414 can be played back via one of the transducers 402, while the other signal 216 is played back by one or more other transducers 402. In another example, the speech content signal 414 can be played back via one or more of the transducers 402, while the other signal 216 is transmitted, via the communication interface 404, to one or more satellite playback devices 408, optionally for playback in synchrony with playback of the speech content signal 414 by the playback device 400. Alternatively, the speech content signal 414 can be transmitted, via the communication interface 404, to another playback device 408 (e.g., a portable playback device 408a, 408b, or 408c) for playback. In some such examples, the playback device 400 may play back the other signal 216, optionally in synchrony with playback of the speech audio signal 404 by the other playback device 408. Numerous variations are envisioned and intended to be part of this disclosure.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0129] In some examples, handling of the speech content signal 414 and the other signal 416 may depend on the playback environment and / or the type of input audio signal 406. For example, if the playback device 400 is in a bonded group with another playback device, the playback device 400 can be configured to transmit the speech content signal 414 to the other playback device. In some examples, this approach can be followed when the other playback device is a particular device (e.g., as may be determined via a device identification) or a particular type of device (e.g., a portable playback device or a headphone device).

[0130] As described above, in some instances, the input audio signal 406 comprises multichannel audio content. For example, referring to FIG. 6, there is illustrated an example in which the input signal 406 comprises multi-channel audio content, including a center channel 602, a left channel 604, a right channel 606, and one or more surround channels 608. In many soundtracks, speech may be solely or predominantly present in the center channel, with little to no speech content present in the left, right, and / or surround channels. Accordingly, in some examples, speech extraction and enhancement processing may be applied only to the center channel, as illustrated in FIG. 6. Thus, in some examples, operation 504 and operations 510 and / or 512 are applied to the center audio channel 602 to produce an enhanced center audio channel signal 610, while the left, right, and surround channels are left unprocessed (from a speech detection / enhancement perspective). In the example of FIG. 6, operations 510 and / or 512 can include relative volume adjustment and / or other processing, as described above.

[0131] As described above, in some examples, the speech extraction operation 504 and separation of the input audio signal 406 into the speech audio 506 and non-speech audio 508 involves a masking technique, as described with reference to FIG. 2. As a result of the masking approach, in some instances, applying the speech enhancement processing at operations 510 and / or 512 can result in some artifacts (e.g., perceptible distortions) being present in the audio output signal 514 due to some non-speech content being present in the same frequency bins as the speech content. Where the input audio signal 406 comprises multi-channel audio content, the problem of such artifacts can be addressed by applying the speech extraction / enhancement processing operations on the center channel only, as in the example of FIG. 6. This approach allows the left and right channels 604, 606 to assist in masking otherwise perceptible artifacts that result from the speech enhancement processing.

[0132] According to certain examples, the speech audio signal 506 produced by operation 504 being applied to one audio channel can be used to control audio processing param eters / settings that are applied to one or more other audio channels of a multi-channel input audio signal 406. An example is illustrated in FIG. 7.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0133] Referring to FIG. 7, in this example, operation 504 is performed on the center audio channel 602, as described above. The resulting speech audio signal 506 is then used to control audio processing that is applied to the center channel at operation 702, to the left channel at operation 704, and to the right channel at operation 706. In some examples, operation 702 corresponds to operation 510 described above. Operations 702, 704, and / or 706 may include adjusting any one or more audio processing parameters, such as equalization settings, gain, dynamic range compression, filtering, transient shaping or other speech feature enhancement, etc., as described above. The processing applied at each of operations 702, 704, and 706 may be the same or different. For example, operation 702 may include increasing the volume of the speech audio 506. Operations 704 and / or 706 may include adjusting equalization settings and / or other processing. As a result of the processing, enhanced / altered center, left, and right channels 708, 710, 712, respectively, may be produced. Such approaches may help to make the speech more present in the center audio channel 610 while also reducing the impact of noise in the left and right channels 710, 712 (e.g., through the use of multiband dynamic processing and / or other audio processing techniques).

[0134] According to certain examples, a combination of techniques described above with reference to FIGS. 6 and 7 can be combined. An example is illustrated in FIG. 8. Thus, as shown in FIG. 8, operation 802 may include applying speech enhancement audio processing (e.g., operations 510 and / or 512 described above) in the center audio channel to produce the enhanced center channel 610, as described above with reference to FIG. 6. In addition, the speech audio signal 506 can be used to control audio processing operations 704 and / or 706 applied to the left and right audio channels 604 and / or 606, respectively, as described above with reference to FIG. 7. This approach may provide a high degree of flexibility to achieve a good balance between overall audio quality and speech intelligibility.

[0135] Although not illustrated in FIGS. 7 or 8, it will be appreciated that in some examples, audio processing based on the speech audio signal 506 can also be applied to one or more of the surround channel(s) 608.

[0136] Further, as described above with reference to FIG. 4A, in some examples, the playback device 400 may transmit any one or more of the audio channels 610, 604, 606, 608, 708, 710, 712, to one or more other playback devices 408 for playback, optionally in synchrony with playback of any one or more of the audio channels by the playback device 400. In some examples, for example, those in which the playback device 400 plays back the center audio channel 610 / 708, the left channel 710 and the right channel 712, the audio processing operations 702, 704, 706, 802, may include adjusting settings associated with components ofAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT the audio processing circuitry 418, as described above. In other examples in which the playback device 400 causes one or more other playback devices 408 to play back any one or more of center audio channel 610 / 708, the left channel 710 and the right channel 712, the audio processing operations 702, 704, 706, 802, may include storing setting information that is then transmitted from the playback device 400 to the one or more other playback devices 408 along with the associated audio channels. Thus, the playback device 400 may instruct the corresponding one or more other playback devices to adjust one or more audio processing settings (e.g., equalization settings, dynamic range compression settings, gain, filtering, transient shaping or other speech feature enhancement settings, etc.) applied for playing back the respective audio channel(s).

[0137] Thus, techniques are provided by which speech audio content can be detected on a dynamic, real-time basis, and when speech is detected speech enhancement audio processing can be applied. These techniques allow for improved speech intelligibility in mixed audio soundtracks, while reducing negative impacts on the overall audio quality or sound experience. As described above, in certain examples, the playback device 400 can leverage machine learning techniques implemented via the MPC 302, for example, to achieve improved performance relative to “always on” speech enhancement approaches. Furthermore, through the MPC 302, speech enhancement processes can be tailored to different playback environments.

[0138] For example, as described above, in some examples, the input data 304 can include information that identifies whether or not the input audio signal 406 comprises multi-channel audio content. Accordingly, based on a determination that the input audio signal 406 does contain multi-channel audio content, the MPC 302 may apply one of the approaches described above with reference to FIGS. 6, 7, or 8, for example. On the other hand, if the input audio signal 406 does not comprise a center audio channel 602, for example, the MPC 302 may operate on the complete input audio signal, such as in the example described above with reference to FIG. 5, for example. In some examples, the approaches described above with reference to FIGS. 6, 7, and / or 8 can be tailored depending on the audio channel content of the input audio signal 406 (e.g., which surround channels are present) and / or the configuration or playback capability of the playback device 400 (e.g., whether or not the playback device 400 includes specific transducers 402 for, center, left, right, and / or other audio channels). Configuration information for the playback device 400) may be context data that is part of the input data 304, or may be programmed into the optimizer 306 for a particular playback device, for example.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0139] Furthermore, the MPC 302 may operate a different approach based on other context data, such as the type of audio content in the input audio signal. For example, as described above, context data in the input data 304 can indicate whether the input audio signal 406 is home theater audio (e.g., a television or movie soundtrack) or music audio, or what genre of music (e.g., opera vs rap or pop) and / or type of video (e.g., an action movie with a potentially high percentage of covered / obscured dialogue due to loud background noise, or a more dialogue-focused drama or documentary) is associated with the input audio signal 406. This context data can be used to adjust various aspects of the speech enhancement process, including, for example, the model 308 selected for use and / or the audio processing applied at any of operations 510, 512, 702, 704, 706, and / or 802. Thus, the system can achieve a high degree of flexibility to adjust to different playback environments.

[0140] In addition, as described above, in certain examples, a user of the playback device 400 can be provided with certain control over the speech enhancement processing to be applied in various circumstances. As described above, user control can be provided at least in part by storing the user preferences 312 that may in turn influence operation of the MPC 302. For example, the user may specify that for certain types of audio content (e.g., music), no speech enhancement processing is to be applied, whereas speech enhancement processing may be applied for home theater audio content. The user may select the type of speech enhancement processing to be applied. For example, the user may specify that only the approach of FIG. 5 is to be applied, or that only the approach of FIG. 6 is to be applied, or some other variation. For example, the user may specify that the approach of FIG. 8 is to be applied for a certain genre of video programming, and that the approach of one of FIGS. 5, 6, or 7 is to be applied for other genres of video programming. In other examples, the user preferences may specify that speech detection / extraction using the machine learning system 300 is not to be applied (at all or in certain circumstances) and that speech enhancement based on estimated known frequency content of human speech, as described above, is to be applied. Numerous other variations are envisioned and form part of this disclosure. Thus, by storing (and updating) the user preferences 312, the user can be provided with a high degree of control over the type and circumstances of speech enhancement used in their MPS 100.IV. Conclusion

[0141] The above discussions relating to playback devices, controller devices, playback zone configurations, and media content sources provide only some examples of operating environments within which functions and methods described below may be implemented. Other operating environments and configurations of media playback systems, playbackAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT devices, and network devices not explicitly described herein may also be applicable and suitable for implementation of the functions and methods.

[0142] The description above discloses, among other things, various example systems, methods, apparatus, and articles of manufacture including, among other components, firmware and / or software executed on hardware. It is understood that such examples are merely illustrative and should not be considered as limiting. For example, it is contemplated that any or all of the firmware, hardware, and / or software aspects or components can be embodied exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Accordingly, the examples provided are not the only ways to implement such systems, methods, apparatus, and / or articles of manufacture.

[0143] Additionally, references herein to “embodiment” means that a particular element, structure, or characteristic described in connection with the embodiment can be included in at least one example embodiment disclosed herein. The appearances of this phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments mutually exclusive of other embodiments. As such, the embodiments described herein, explicitly and implicitly understood by one skilled in the art, can be combined with other embodiments.

[0144] The specification is presented largely in terms of illustrative environments, systems, procedures, steps, logic blocks, processing, and other symbolic representations that directly or indirectly resemble the operations of data processing devices coupled to networks. These process descriptions and representations are typically used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art. Numerous specific details are set forth to provide a thorough understanding of the present disclosure. However, it is understood to those skilled in the art that certain embodiments of the present disclosure can be practiced without certain, specific details. In other instances, well known methods, procedures, components, and circuitry have not been described in detail to avoid unnecessarily obscuring aspects of the embodiments. Accordingly, the scope of the present disclosure is defined by the appended claims rather than the foregoing description of embodiments.

[0145] When any of the appended claims are read to cover a purely software and / or firmware implementation, at least one of the elements in at least one example is hereby expressly defined to include a tangible, non-transitory medium such as a memory, DVD, CD, Blu-ray, and so on, storing the software and / or firmware.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCTV. Additional Examples

[0146] The following examples pertain to further embodiments, from which numerous permutations and configurations will be apparent.

[0147] Example 1 provides a method comprising detecting, with a playback device, an audio signal; applying a parametric machine learning model to dynamically detect speech in the audio signal; based on detecting the speech, separating the audio signal into speech audio and nonspeech audio; applying first audio processing to the speech audio to produce processed speech audio; applying second audio processing to the non-speech audio to produced processed nonspeech audio, the second audio processing being different from the first audio processing; combining the processed speech audio and the processed non-speech audio to produce an audio output signal; and playing back the audio output signal via the playback device.

[0148] Example 2 includes the method of Example 1, wherein applying the first audio processing comprises increasing a volume of the speech audio, and wherein applying the second audio processing comprises decreasing a volume of the non-speech audio.

[0149] Example 3 includes the method of one of Examples 1 or 2, further comprising performing a short-time Fourier transform on the audio signal to produce a first signal, the first signal being a frequency domain representation of the audio signal, wherein applying the parametric machine learning model includes processing the first signal with the parametric machine learning model.

[0150] Example 4 includes the method of Example 3, wherein separating the audio signal into the speech audio and the non-speech audio comprises identifying first frequency content of the first signal corresponding to the speech audio.

[0151] Example 5 includes the method of any one of Examples 1-4, wherein applying the parametric machine learning model comprises processing the audio signal with a deep neural network trained to detect speech.

[0152] Example 6 includes the method of any one of Examples 1-5, wherein the audio signal comprises multi-channel audio content including a plurality of audio channels, wherein the plurality of audio channels comprises a center audio channel and at least one other audio channel, and wherein applying a parametric machine learning model to dynamically detect speech in the audio signal comprises applying the parametric machine learning model to the center audio channel to detect the speech in the center audio channel.

[0153] Example 7 includes the method of Example 6, wherein combining the processed speech audio and the processed non-speech audio to produce an audio output signal comprises combining the processed speech audio and the processed non-speech audio to produce aAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT processed center audio channel, and wherein playing back the audio output signal comprises playing back the processed center audio channel and the at least one other audio channel.

[0154] Example 8 includes the method of Example 7, wherein applying the first audio processing comprises increasing a volume of the speech audio in the center audio channel, and wherein applying the second audio processing comprises decreasing a volume of the nonspeech audio in the center audio channel.

[0155] Example 9 includes the method of Example 6, wherein combining the processed speech audio and the processed non-speech audio to produce an audio output signal comprises combining the processed speech audio and the processed non-speech audio to produce a processed center audio channel; and wherein playing back the audio output signal comprises: playing back the processed center audio channel; and causing at least one other playback device to play back the at least one other audio channel.

[0156] Example 10 includes the method of Example 6, wherein the playback device comprises a plurality of audio transducers, wherein applying the first audio processing comprises applying a first set of equalization settings for at least one audio transducer configured to play back the center audio channel, and wherein applying the second audio processing comprises applying a second set of equalization settings for at least one other audio transducer configured to play back the at least one other audio channel.

[0157] Example 11 includes the method of Example 6, further comprising: based on detecting the speech in the center audio channel, applying third audio processing to the at least one other audio channel.

[0158] Example 12 provides a playback device configured to implement the method of any one of Examples 1-11.

[0159] Example 13 provides a playback device comprising: a communication interface configured to detect an audio signal, wherein the audio signal comprises speech audio and non- speech audio; a plurality of audio transducers; at least one processor; and at least one non- transitory computer-readable medium storing program instructions that are executable by the at least one processor to cause the playback device to apply a parametric machine learning model to the audio signal to detect the speech audio in the audio signal, based on detecting the speech audio, process the audio signal to increase a volume of the speech audio and decrease a volume of the non-speech audio, thereby producing a processed audio signal, and play back the processed audio signal via the plurality of audio transducers.

[0160] Example 14 includes the playback device of Example 13, wherein the parametric machine learning model is a deep neural network trained to detect human speech.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0161] Example 15 includes the playback device of one of Examples 13 or 14, wherein the at least one non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to perform a short-time Fourier transform on the audio signal to produce a first signal, the first signal being a frequency domain representation of the audio signal.

[0162] Example 16 includes the playback device of Example 15, wherein the at least one non- transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to apply the parametric machine learning model to the first signal to produce, based on detecting the speech audio, a speech mask that identifies first frequency content of the first signal corresponding to the speech audio.

[0163] Example 17 includes the playback device of Example 16, wherein the at least one non- transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to multiply the first signal by the speech mask to produce a processed speech signal.

[0164] Example 18 includes the playback device of Example 17, wherein the at least one non- transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to invert the speech mask to produce a noise mask, and multiply the first signal by the noise mask to produce a processed non-speech signal.

[0165] Example 19 includes the playback device of Example 18, wherein the at least one non- transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to combine the processed speech signal and the processed non-speech signal to produce a second signal in a frequency domain, and perform an inverse short-time Fourier transform on the second signal to produce the processed audio signal.

[0166] Example 20 includes the playback device of Example 15, wherein to perform the short- time Fourier transform on the audio signal, the at least one non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to perform the short-time Fourier transform to produce the first signal represented by a number of frequency bins, wherein the number of frequency bins is configurable based on one or more attributes of the audio signal.

[0167] Example 21 includes the playback device of Example 20, wherein the one or more attributes of the audio signal include an acoustic genre of the audio signal and / or a codec of the audio signal.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0168] Example 22 includes the playback device of any one of Examples 13-21, wherein the audio signal is a center audio channel of a multi-channel audio signal that comprises the center audio channel and one or more other audio channels.

[0169] Example 23 includes the playback device of Example 22, wherein the at least one non- transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to play back, via the plurality of audio transducers, at least one audio channel of the one or more other audio channels.

[0170] Example 24 includes the playback device of Example 22, wherein the at least one non- transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to cause another playback device to play back at least one audio channel of the one or more other audio channels.

[0171] Example 25 includes the playback device of one of Examples 23 or 24, wherein the playback device is a soundbar.

[0172] Example 26 includes the playback device of any one of Examples 22-25, wherein the at least one non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to adjust audio signal processing applied to at least one of the one or more other audio channels based on the speech audio.

[0173] Example 27 provides a method comprising: detecting an audio signal comprising multichannel audio content including a center audio channel and a plurality of additional audio channels; applying a parametric machine learning model to the center audio channel to detect speech in the center audio channel; and based on detecting the speech, controlling audio processing applied to the center audio channel and to at least one additional audio channel of the plurality of additional audio channels.

[0174] Example 28 includes the method of Example 27, wherein the multi-channel audio content is home theater audio content, and wherein the plurality of additional audio channels includes a front left channel, a front right channel, and one or more surround channels.

[0175] Example 29 includes the method of one of Examples 27 or 28, wherein controlling the audio processing comprises: applying, based on detecting the speech, first audio processing to the center audio channel; and applying, based on detecting the speech, second audio processing to the front left channel and to the front right channel.

[0176] Example 30 includes the method of Example 29, wherein applying the parametric machine learning model to the center audio channel to detect speech in the center audio channel comprises identifying speech audio content in the center audio channel and non-speech audioAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT content in the center audio channel, and wherein applying the first audio processing to the center audio channel includes increasing a volume of the speech audio content relative to the non-speech audio content.

[0177] Example 31 includes the method of any one of Examples 27-30, further comprising performing a short-time Fourier transform on the center audio channel to produce a first signal, the first signal being a frequency domain representation of the center audio channel, wherein applying a parametric machine learning model to the center audio channel comprises identifying first frequency content in the first signal corresponding to the speech.

[0178] Example 32 includes the method of Example 31, wherein applying the parametric machine learning model comprises producing, based on the first frequency content, a speech mask.

[0179] Example 33 includes the method of Example 32, wherein controlling audio processing applied to the center audio channel comprises: multiplying the first signal by the speech mask to produce a processed speech signal; inverting the speech mask to produce a noise mask; multiplying the first signal by the noise mask to produce a processed non-speech signal; and combining the processed speech signal and the processed non-speech signal to produce a second signal.

[0180] Example 34 includes the method of Example 33, further comprising performing an inverse short-time Fourier transform on the second signal to produce a time-domain signal representing a processed version of the center audio channel.

[0181] Example 35 includes the method of any one of Examples 31-34, further comprising adjusting a frequency sample size of the short-time Fourier transform based on one of more attributes of the audio signal.

[0182] Example 36 provides a playback device configured to implement the method of any one of Examples 27-35.

[0183] Example 37 provides a playback device comprising: a communication interface configured to detect an audio signal, wherein the audio signal comprises multi-channel audio content including a center audio channel and a plurality of additional audio channels; a plurality of audio transducers; at least one processor; and a non-transitory computer-readable medium storing program instructions that are executable by the at least one processor to cause the playback device to apply a parametric machine learning model to the center audio channel to detect speech in the center audio channel, based on detecting the speech, (i) process the center audio channel to produce a processed center audio channel, and (ii) control audio processing applied to at least one additional audio channel of the plurality of additional audio channels,Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT play back, via the plurality of transducers, the processed center audio channel, and cause playback of the plurality of additional audio channels.

[0184] Example 38 includes the playback device of Example 37, wherein the playback device is a soundbar.

[0185] Example 39 includes the playback device of one of Examples 37 or 38, wherein the multi-channel audio content is home theater audio content, and wherein the plurality of additional audio channels includes a front left channel, a front right channel, and one or more surround channels.

[0186] Example 40 includes the playback device of Example 39, wherein to cause playback of the plurality of additional audio channels, the non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to transmit, via the communication interface, the at least one additional audio channel to a corresponding at least one additional playback device.

[0187] Example 41 includes the playback device of Example 40, wherein to control the audio processing applied to the at least one additional audio channel, the non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to instruct the corresponding at least one additional playback device to adjust one or more equalization settings applied for playing back the at least one additional audio channel.

[0188] Example 42 includes the playback device of any one of Examples 37-41, wherein to process the center audio channel, the non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to increase a volume of speech audio content relative to non-speech audio content in the center audio channel.

[0189] Example 43 includes the playback device of Example 42, wherein the non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to perform a short-time Fourier transform on the center audio channel to produce a first signal, the first signal being a frequency domain representation of the center audio channel; and wherein to apply the parametric machine learning model to the center audio channel, the at least one non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to apply the parametric machine learning model to the first signal to: identify first frequency content in the first signal corresponding to the speech, and produce a speech mask based on the first frequency content.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0190] Example 44 includes the playback device of Example 43, wherein to control audio processing applied to the center audio channel, the non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to multiply the first signal by the speech mask to produce a processed speech signal; invert the speech mask to produce a noise mask; multiply the first signal by the noise mask to produce a processed non-speech signal; combine the processed speech signal and the processed non-speech signal to produce a second signal; and perform an inverse short-time Fourier transform on the second signal to produce the processed center audio channel.

[0191] Example 45 provides a playback device comprising: a communication interface configured to detect an audio signal comprising multi-channel audio content including a plurality of audio channels; a plurality of audio transducers; at least one processor; and a non- transitory computer-readable medium storing program instructions that are executable by the at least one processor to cause the playback device to apply a short-time Fourier transform to at least a portion of the audio signal to produce a first signal, the first signal being a frequency domain representation of the audio signal, based on a user-configurable control setting, process the first signal to identify first frequency content associated with speech in the audio signal; based on the first frequency content, control audio processing applied to at least one audio channel of the plurality of audio channels, and play back, via the plurality of audio transducers, one or more audio channels of the plurality of audio channels, wherein to process the first signal comprises (i) to apply, based on the user-configuration control setting having a first value, a parametric machine learning model to the first signal to detect the speech and identify the first frequency content based on detecting the speech, or (ii) to identify, based on the userconfiguration control setting having a second value, the first frequency content by selecting plurality of pre-set frequency bins.

[0192] Example 46 provides a method comprising: detecting, with a playback device, an audio signal; producing, based at least in part on the audio signal, input data comprising audio data and context data, the audio data representing audio content of the audio signal and the context data describing one or more attributes of the audio signal and / or a playback environment in which the audio signal is to be played; based at least in part on the context data, applying a speech-based audio processing process including (i) applying a parametric machine learning model to at least a portion of the audio data to detect speech content in the audio signal, (ii) based on detecting the speech, extracting a speech signal from audio signal, and (iii) using the speech signal, controlling audio processing of at least a portion of the audio signal to produceAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT a processed audio signal; and causing playback of at least a portion of the processed audio signal.

[0193] Example 47 includes the method of Example 46, wherein the one or more attributes of the audio signal include an acoustic genre of the audio signal and / or a codec of the audio signal.

[0194] Example 48 includes the method of one of Examples 46 or 47, wherein the context data includes information specifying that the playback environment comprises a bonded group including the playback device and at least one additional playback device.

[0195] Example 49 includes the method of Example 48, wherein the audio content comprises multi-channel audio content, and wherein applying the parametric machine learning model to at least the portion of the audio data comprises applying the parametric machines learning model to a portion of the audio data representing a center audio channel; and wherein extracting the speech signal comprises extracting the speech signal from the center audio channel of the audio signal.

[0196] Example 50 includes the method of Example 49, wherein controlling audio processing of at least a portion of the audio signal comprises: controlling audio processing of the center audio channel to produce a processed center audio channel, and controlling audio processing of the at least one other audio channel to produce a processed at least one other audio channel.

[0197] Example 51 includes the method of Example 50, wherein the at least one other audio channel includes a left channel and a right channel.

[0198] Example 52 includes the method of one of Examples 50 or 51, wherein causing playback of at least a portion of the processed audio signal comprises transmitting the processed at least one other audio channel to the at least one additional playback device.

[0199] Example 53 includes the method of Example 52, wherein the processed at least one other audio channel comprises: audio content of the audio signal corresponding to the at least one other audio channel; and audio processing control information identifying one or more audio processing settings to be applied by the at least one additional playback device for playback of the at least one other audio channel.

[0200] Example 54 includes the method of one of Examples 52 or 53, wherein causing playback of at least a portion of the processed audio signal comprises playing back, with the playback device, the processed center audio channel.

[0201] Example 55 includes the method of any one of Examples 46-54, wherein applying the parametric machine learning model comprises processing the portion of the audio data with a deep neural network trained to detect speech.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0202] Example 56 includes a media playback system comprising one or more playback devices configured to perform the method of any one of Examples 46-55.

[0203] Example 57 provides a method comprising: detecting, with a playback device, an audio signal; producing, based at least in part on the audio signal, input data comprising audio data representing audio content in the audio signal; applying a parametric machine learning model to at least a portion of the audio data to detect speech content in the audio signal; based on detecting the speech, extracting a speech signal from audio signal; using the speech signal, applying a speech-based audio processing process to at least a portion of the audio signal to produce a processed audio signal; and causing playback of at least a portion of the processed audio signal.

[0204] Example 58 includes the method of Example 57, wherein applying the parametric machine learning model comprises processing the portion of the audio data with a deep neural network trained to detect human speech.

[0205] Example 59 includes the method of one of Examples 57 or 58, wherein the audio content comprises multi-channel audio content including a center audio channel and at least one other audio channel; wherein applying the parametric machine learning model to at least the portion of the audio data comprises applying the parametric machine learning model to a portion of the audio data representing the center audio channel; and wherein extracting the speech signal comprises extracting the speech signal from the center audio channel of the audio signal.

[0206] Example 60 includes the method of Example 59, wherein applying the speech-based audio processing process to at least the portion of the audio signal to produce the processed audio signal comprises: controlling audio processing of the center audio channel to produce a processed center audio channel; and controlling audio processing of the at least one other audio channel to produce a processed at least one other audio channel.

[0207] Example 61 includes the method of Example 60, wherein causing playback of at least the portion of the processed audio signal comprises transmitting the processed at least one other audio channel to at least one other playback device.

[0208] Example 62 includes the method of Example 61, wherein causing playback of at least the portion of the processed audio signal comprises playing back, with the playback device, the processed center audio channel.

[0209] Example 63 includes the method of any one of Examples 59-62, wherein the at least one other audio channel includes a left channel and a right channel.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT

[0210] Example 64 includes the method of one of Examples 57 or 58, wherein the audio content comprises multi-channel audio content including a center audio channel and a plurality of other audio channels; wherein applying the parametric machine learning model to at least the portion of the audio data comprises applying the parametric machine learning model to a portion of the audio data representing the center audio channel and one or more of the plurality of other audio channels; and wherein extracting the speech signal comprises extracting the speech signal from the center audio channel of the audio signal.

[0211] Example 65 provides the playback device comprising: a communication interface configured; one or more audio transducers; at least one processor; and a non-transitory computer-readable medium storing program instructions that are executable by the at least one processor to cause the playback device to perform the method of any one of Examples 57-64.

[0212] Example 66 provides a method comprising: detecting, with a playback device, an audio signal; producing, based at least in part on the audio signal, input data comprising audio data and context data, the audio data representing audio content of the audio signal and the context data describing one or more attributes of the audio signal and / or a playback environment in which the audio signal is to be played; applying a parametric machine learning model to at least a portion of the audio data to detect speech content in the audio signal; based on detecting the speech, extracting a speech signal from the audio signal; based at least in part on the context data, transmitting the speech signal to another playback device; and causing playback of the speech signal by the other playback device.

[0213] Example 67 includes the method of Example 66, wherein the other playback device is a portable playback device.

[0214] Example 68 includes the method of one of Examples 66 or 67, wherein the one or more attributes of the audio signal include an acoustic genre of the audio signal and / or a codec of the audio signal.

[0215] Example 69 includes the method of any one of Examples 66-68, wherein the context data includes information specifying that the playback environment comprises a bonded group including the playback device and the other playback device.

[0216] Example 70 includes the method of any one of Examples 66-69, wherein applying the parametric machine learning model comprises processing the portion of the audio data with a deep neural network trained to detect speech.

[0217] Example 71 provides a media playback system comprising the playback device and the other playback device, wherein the media playback system is configured to implement the method of any one of Examples 66-70.

Claims

Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCTCLAIMS1. A method comprising: detecting, with a playback device, an audio signal; applying a parametric machine learning model to dynamically detect speech in the audio signal; based on detecting the speech, separating the audio signal into speech audio and nonspeech audio; applying first audio processing to the speech audio to produce processed speech audio; applying second audio processing to the non-speech audio to produced processed nonspeech audio, the second audio processing being different from the first audio processing; combining the processed speech audio and the processed non-speech audio to produce an audio output signal; and playing back the audio output signal via the playback device.

2. The method of claim 1, wherein applying the first audio processing comprises increasing a volume of the speech audio; and wherein applying the second audio processing comprises decreasing a volume of the non-speech audio.

3. The method of one of claims 1 or 2, further comprising: performing a short-time Fourier transform on the audio signal to produce a first signal, the first signal being a frequency domain representation of the audio signal; wherein applying the parametric machine learning model includes processing the first signal with the parametric machine learning model.

4. The method of claim 3, wherein separating the audio signal into the speech audio and the non-speech audio comprises identifying first frequency content of the first signal corresponding to the speech audio.

5. The method of any one of claims 1-4, wherein applying the parametric machine learning model comprises processing the audio signal with a deep neural network trained to detect speech.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT6. The method of any one of claims 1-5, wherein the audio signal comprises multichannel audio content including a plurality of audio channels, wherein the plurality of audio channels comprises a center audio channel and at least one other audio channel; and wherein applying a parametric machine learning model to dynamically detect speech in the audio signal comprises applying the parametric machine learning model to the center audio channel to detect the speech in the center audio channel.

7. The method of claim 6, wherein combining the processed speech audio and the processed non-speech audio to produce an audio output signal comprises combining the processed speech audio and the processed non-speech audio to produce a processed center audio channel; and wherein playing back the audio output signal comprises playing back the processed center audio channel and the at least one other audio channel.

8. The method of claim 6, wherein combining the processed speech audio and the processed non-speech audio to produce an audio output signal comprises combining the processed speech audio and the processed non-speech audio to produce a processed center audio channel; and wherein playing back the audio output signal comprises: playing back the processed center audio channel; and causing at least one other playback device to play back the at least one other audio channel.

9. The method of claim 6, wherein the playback device comprises a plurality of audio transducers; wherein applying the first audio processing comprises applying a first set of equalization settings for at least one audio transducer configured to play back the center audio channel; and wherein applying the second audio processing comprises applying a second set of equalization settings for at least one other audio transducer configured to play back the at least one other audio channel.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT10. The method of any one of claims 6-9, further comprising: based on detecting the speech in the center audio channel, applying third audio processing to the at least one other audio channel, the third audio processing being different from the first audio processing and the second audio processing.

11. A playback device comprising: a communication interface configured to detect an audio signal, wherein the audio signal comprises speech audio and non-speech audio; a plurality of audio transducers; at least one processor; and at least one non-transitory computer-readable medium storing program instructions that are executable by the at least one processor to cause the playback device to apply a parametric machine learning model to the audio signal to detect the speech audio in the audio signal, based on detecting the speech audio, process the audio signal to increase a volume of the speech audio and decrease a volume of the non-speech audio, thereby producing a processed audio signal, and play back the processed audio signal via the plurality of audio transducers.

12. The playback device of claim 11, wherein the parametric machine learning model is a deep neural network trained to detect human speech.

13. The playback device of one of claims 11 or 12, wherein the at least one non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to: perform a short-time Fourier transform on the audio signal to produce a first signal, the first signal being a frequency domain representation of the audio signal; and apply the parametric machine learning model to the first signal to produce, based on detecting the speech audio, a speech mask that identifies first frequency content of the first signal corresponding to the speech audio.

14. The playback device of claim 13, wherein the at least one non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to:Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT multiply the first signal by the speech mask to produce a processed speech signal, invert the speech mask to produce a noise mask; multiply the first signal by the noise mask to produce a processed non-speech signal. combine the processed speech signal and the processed non-speech signal to produce a second signal in a frequency domain; and perform an inverse short-time Fourier transform on the second signal to produce the processed audio signal.

15. The playback device of one of claims 13 or 14, wherein to perform the short-time Fourier transform on the audio signal, the at least one non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to: perform the short-time Fourier transform to produce the first signal represented by a number of frequency bins; wherein the number of frequency bins is configurable based on one or more attributes of the audio signal.

16. The playback device of any one of claims 11-15, wherein the audio signal is a center audio channel of a multi-channel audio signal that comprises the center audio channel and one or more other audio channels; and wherein the at least one non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to adjust audio signal processing applied to at least one of the one or more other audio channels based on the speech audio.

17. A playback device comprising: a communication interface configured to detect an audio signal, wherein the audio signal comprises multi-channel audio content including a center audio channel and a plurality of additional audio channels; a plurality of audio transducers; at least one processor; and a non-transitory computer-readable medium storing program instructions that are executable by the at least one processor to cause the playback device toAttorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT apply a parametric machine learning model to the center audio channel to detect speech in the center audio channel, based on detecting the speech,(i) process the center audio channel to produce increase a volume of speech audio content relative to non-speech audio content in the center audio channel to produce a processed center audio channel, and(ii) control audio processing applied to at least one additional audio channel of the plurality of additional audio channels, play back, via the plurality of transducers, the processed center audio channel, and cause playback of the plurality of additional audio channels.

18. The playback device of claim 17, wherein the multi-channel audio content is home theater audio content and the plurality of additional audio channels includes a front left channel, a front right channel, and one or more surround channels; and wherein to cause playback of the plurality of additional audio channels, the non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to: transmit, via the communication interface, the at least one additional audio channel to a corresponding at least one additional playback device.

19. The playback device of claim 40, wherein the non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to perform a short-time Fourier transform on the center audio channel to produce a first signal, the first signal being a frequency domain representation of the center audio channel; and wherein to apply the parametric machine learning model to the center audio channel, the at least one non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to apply the parametric machine learning model to the first signal to: identify first frequency content in the first signal corresponding to the speech, and produce a speech mask based on the first frequency content.Attorney Docket No. SON00087WOU1Client Docket No. 23-1205-PCT20. The playback device of claim 41, wherein to control audio processing applied to the center audio channel, the non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to multiply the first signal by the speech mask to produce a processed speech signal; invert the speech mask to produce a noise mask; multiply the first signal by the noise mask to produce a processed non-speech signal; combine the processed speech signal and the processed non-speech signal to produce a second signal; and perform an inverse short-time Fourier transform on the second signal to produce the processed center audio channel.

Citation Information

Patent Citations

  • Voice control of a media playback system

    US10499146B2

  • Room association based on name

    US10712997B2

  • System and method for synchronizing operations among a plurality of independently clocked digital data processing devices

    US8234395B2

  • Controlling and manipulating groupings in a multi-zone media system

    US8483853B1

  • Method and apparatus for processing an initial audio signal

    US20230087486A1