Calibrating playback entities in playback networks comprising multiple playback entities

By calibrating playback settings through boundary condition determination, spectral and spatial analysis, and low frequency analysis, the method addresses the challenges of acoustic reflections and low frequency content in large listening areas, improving audio quality in playback systems.

WO2025111229A1PCT designated stage expired Publication Date: 2025-05-30SONOS INC
View PDF 45 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/056425
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-11-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing playback systems face challenges in calibrating multiple playback entities in large listening areas with distributed speakers, leading to issues with acoustic reflections and coherent addition of low frequency content, resulting in suboptimal audio playback.

Method used

The implementation of a method for calibrating playback settings in Area Zones, which involves determining boundary conditions, performing spectral and spatial analysis, and conducting low frequency analysis to adjust equalization settings and mitigate acoustic effects.

Benefits of technology

This approach enhances audio quality by reducing the impact of acoustic reflections and excessive bass levels, providing a more balanced and coherent audio experience in large listening areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024056425_30052025_PF_FP_ABST
    Figure US2024056425_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed embodiments include systems and methods for calibrating playback entities in grouped playback configurations, including area zones configurations. In some embodiments, calibrating the playback entities includes (i) determining one or more boundary conditions for each playback entity in the group, (ii) conducting one or more spectral and / or spatial analyses and calibrations for the playback entities in the playback group, and (iii) performing a low frequency analysis and correction for the playback group, and in some instances, performing a low frequency analysis and correction for two or more groups together in scenarios where the two or more groups are implemented within the same listening area.
Need to check novelty before this filing date? Find Prior Art

Description

CALIBRATING PLAYBACK ENTITIES IN PLAYBACK NETWORKS COMPRISING MULTIPLE PLAYBACK ENTITIESCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional App. 63 / 601,155, titled “Calibrating Playback Entities in Playback Networks Comprising Multiple Playback Entities,” filed on Nov. 20, 2023, and currently pending. The entire contents of U.S. Provisional App. 63 / 601,155 are incorporated herein by reference.

[0002] This application is also related to and incorporates by reference the entire contents of the following applications: (i) U.S. Provisional App. 63 / 377,948, titled “Playback System Architecture,” filed on Sep. 30, 2022, and now expired; (ii) U.S. Provisional App. 63 / 377,899, titled “Multichannel Content Distribution,” filed on Sep. 30 2022, and now expired; (iii) U.S. Provisional App. 63 / 377,967, titled “Playback Systems with Dynamic Forward Error Correction,” filed on Sep. 30, 2022, and now expired; (iv) U.S. Provisional App. 63 / 377,978, titled “Broker / Sub scriber Model for Information Sharing and Management Among Connected Devices,” filed on Sep. 30, 2022, and now expired; (v) U.S. Provisional App. 63 / 377,979, titled “Multiple Broker Deployment for Information Sharing and Management Among Connected Devices,” filed on Sep. 30, 2022, and now expired; (vi) U.S. Provisional App. 63 / 377,978, titled “Global State Service,” filed on Sep. 30, 2022, and now expired; (vi) U.S. Provisional App. 63 / 502,347, titled “Area Zones,” filed on May 15, 2023, and now expired; (vii) U.S. App. 18 / 478,063, titled “Multichannel Content Distribution,” filed on Nov. 29, 2023, and currently pending; (viii) PCT App. PCT / US2023 / 034170 titled “Playback System Architectures and Area Zone Configurations,” filed on Sep. 29, 2023, and published on Apr. 4, 2024, as Inf 1 Pub. WO 2024 / 073078; and (ix) PCT App. PCT / US2023 / 034181 titled “State information exchange among connected devices,” filed on Sep. 29, 2023, and published on Apr. 4, 2024, as Int’l Pub. WO 2024 / 073086.

[0003] Aspects of the features and functions disclosed and described in the aboveidentified applications can be used in combination with the examples disclosed and described herein and with each other in some instances to improve the functionality and performance of playback systems including playback systems having large numbers of playback entities, including but not limited to playback systems having playback entities implemented in Area Zone configurations.FIELD OF THE DISCLOSURE

[0004] The present disclosure is related to consumer goods and, in some more particular examples, to methods, systems, products, features, services, and other elements relating to media playback systems, media playback devices, methods of operating media playback systems and devices, and various features and aspects thereof.BACKGROUND

[0005] Options for accessing and listening to digital audio in an out-loud setting were limited until in 2002, when SONOS, Inc. began development of a new type of playback system. Sonos then fded one of its first patent applications in 2003, titled “Method for Synchronizing Audio Playback between Multiple Networked Devices,” and began offering its first media playback systems for sale in 2005. The Sonos Wireless Home Sound System enables people to experience music from many sources via one or more networked playback devices. Through a software control application installed on a controller (e.g., smartphone, tablet, computer, voice input device), individuals can play most any music they like in any room having a networked playback device. Media content (e.g., songs, podcasts, video sound) can be streamed to playback devices such that each room with a playback device can play back corresponding different media content. In addition, rooms can be grouped together for synchronous playback of the same media content, and / or the same media content can be heard in all rooms synchronously.BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Features, aspects, and advantages of the presently disclosed technology may be better understood with regard to the following description, appended claims, and accompanying drawings, as listed below. A person skilled in the relevant art will understand that the features shown in the drawings are for purposes of illustrations, and variations, including different and / or additional features and arrangements thereof, are possible.

[0007] Figure 1A shows a partial cutaway view of an environment having a media playback system configured in accordance with aspects of the disclosed technology.

[0008] Figure IB shows a schematic diagram of the media playback system of Figure 1A and one or more networks.

[0009] Figure 1C shows a block diagram of a playback device.

[0010] Figure ID shows a block diagram of a playback device.

[0011] Figure IE shows a block diagram of a network microphone device.

[0012] Figure IF shows a block diagram of a network microphone device.

[0013] Figure 1G shows a block diagram of a playback device.

[0014] Figure 1H shows a partially schematic diagram of a control device.

[0015] Figures 1-1 through IL show schematic diagrams of corresponding media playback system zones.

[0016] Figure IM shows a schematic diagram of media playback system areas.

[0017] Figure 2A shows a front isometric view of a playback device configured in accordance with aspects of the disclosed technology.

[0018] Figure 2B shows a front isometric view of the playback device of Figure 3 A without a grille.

[0019] Figure 2C shows an exploded view of the playback device of Figure 2A.

[0020] Figure 3A shows a front view of a network microphone device configured in accordance with aspects of the disclosed technology.

[0021] Figure 3B shows a side isometric view of the network microphone device of Figure 3A.

[0022] Figure 3C shows an exploded view of the network microphone device of Figures 3A and 3B.

[0023] Figure 3D shows an enlarged view of a portion of Figure 3B.

[0024] Figure 3E shows a block diagram of the network microphone device of Figures3A-3D

[0025] Figure 3F shows a schematic diagram of an example voice input.

[0026] Figures 4A-4D show schematic diagrams of a control device in various stages of operation in accordance with aspects of the disclosed technology.

[0027] Figure 5 shows front view of a control device.

[0028] Figure 6 shows a message flow diagram of a media playback system.

[0029] Figure 7A shows a network of playback entities configured into a single AreaZone within a single room according to some example embodiments.

[0030] Figure 7B shows a network of playback entities configured into a two Area Zones within the same room according to some example embodiments.

[0031] Figure 7C shows a network of playback entities configured into two Area Zones, where one of the Area Zones includes playback entities in two separate rooms, according to some example embodiments.

[0032] Figure 7D shows a network of playback entities configured into two Area Zones in two rooms, where the first Area Zone includes playback entities configured to play audio in a first room, where the second Area Zone includes playback entities configured to play audio in the first room and a second room, where the playback entities in the second Area Zone include a Multi-Player Playback Device, and where the Multi-Player Playback Device comprises a first logical player configured to play audio in the first room and a second logical player configured to play audio in the second room, according to some example embodiments.

[0033] Figure 7E shows a network of playback entities configured into two Area Zones in two rooms implemented via two Multi-Player Playback Devices, where the first Multi-Player Playback Device comprises (i) a first logical player configured to play audio in a first Area Zone within a first room, where the first Area Zone includes the first logical player of the first Multi- Player Playback Device and a separate logical player implemented by a second Multi-Player Playback Device, and (ii) a second logical player configured to play audio in a second Area Zone within a second room, according to some example embodiments.

[0034] Figure 8 shows an example Multi-Player Playback Device configurable to implement from one to eight logical players according to some example embodiments.

[0035] Figure 9A shows steps of an example method depicting aspects of calibrating playback entities in playback networks comprising multiple playback entities according to some example embodiments.

[0036] Figure 9B shows further steps of an example method depicting aspects of calibrating playback entities in playback networks comprising multiple playback entities according to some example embodiments.

[0037] Figure 9C shows still further steps of an example method depicting aspects of calibrating playback entities in playback networks comprising multiple playback entities according to some example embodiments

[0038] The drawings are for the purpose of illustrating example configurations, but those of ordinary skill in the art will understand that the technology disclosed herein is not limited to the arrangements and / or instrumentality shown in the drawings.DETAILED DESCRIPTIONI. Overview

[0039] Existing playback devices, examples of which include Sonos’ s intelligent playback devices, can be used to implement different types of groupings of playback devices configured to play audio content together with each other in a groupwise fashion. Playback groupings implemented with intelligent playback devices such as the playback devices available from Sonos are very flexible in terms of configuration options and the ability to play many different types of audio content, thereby enabling many different groupwise configurations for listening to many different types of audio content in many different types of listening environments.

[0040] Embodiments disclosed herein provide enhanced flexibility and scalability over existing types of grouped configurations via the introduction of several new concepts, examples of which include (i) a new type of grouping referred to herein as an “Area Zone” which implements a new type of hierarchical control plane to streamline signaling between and among playback entities within the Area Zone to support audio distribution to and playback by a relatively large number of playback entities within the Area Zone, (ii) a new type of playback device referred to herein as a Multi-Player Playback Device that includes one or more processors, tangible, non-transitory computer readable media and numerous (such as, in an example, up to eight) configurable audio outputs, where one example of such a new Multi-Player Playback Device is configurable to implement from one to eight logical “players,” and (iii) a new type of logical “player” implemented by a Multi-Player Playback Device. These and other aspects of this new “Area Zone” grouping and these new physical and logical playback entities are described herein. Some aspects of Area Zones and Area Zone configurations containing playback entities are described herein. Further aspects of Area Zones are disclosed and described in U.S. Provisional App. 63 / 502,347, titled “Area Zones,” referred to as Docket No. 22-1002p, filed on May 15, 2023, the entire contents of which are incorporated herein by reference.

[0041] This new “Area Zone” configuration presents new technical challenges in scenarios where a listening area has larger numbers of playback entities, particularly where the speakers associated with the playback entities are distributed throughout the listening area. Listening areas with larger numbers of playback entities are contemplated in commercial spaces,such as retail stores, restaurants, hotels, bars, theme parks, and other locations. However, Area Zones and the features associated therewith are equally applicable to, and can have similar advantages for, listening areas having smaller groups of playback entities as well.

[0042] Some contemplated Area Zone configurations may include playback entities having speakers positioned in or near ceilings, corners, cabinets, shelves, or other acoustically reflective surfaces, boundaries, or barriers that may affect how the audio played via the Area Zone within a room (or other listening area) is perceived by listeners, particularly in Area Zone configurations having several speakers positioned near reflective surfaces. Similarly, Area Zones where several playback entities are driving speakers within the same room may result in coherent addition of low frequency content, leading to excessive bass levels within the listening area. The effects of acoustic reflections and coherent addition of low frequency content often tend to become more pronounced in Area Zone configurations having greater numbers of playback entities compared to a typical household, for instance.

[0043] To overcome or at least ameliorate at least some of the above-described acoustic effects resulting from both speakers positioned near acoustically reflective surfaces and coherent addition of low frequency content, some embodiments include methods and processes for calibrating playback settings for the playback entities of an Area Zone within a listening area. One goal of calibrating the playback settings for the playback entities in this manner is to adjust equalization settings that the playback entities use when playing audio within the listening area so as to offset or otherwise pre-compensate for the effects of the acoustic reflections and the coherent addition of low frequency content.

[0044] For example, to offset or pre-compensate for the effects of acoustic reflections, playback entities having speakers positioned near acoustically reflective surfaces can be identified (e.g., via boundary condition determination methods disclosed herein). And the playback entities having speakers positioned near acoustically reflective surfaces can be configured with equalization settings that reduce the gain applied to (or perhaps apply negative to) acoustic frequencies that are reflected by those acoustically reflective surfaces so as to reduce the extent to which the audio heard by listeners is affected by the acoustic reflections.

[0045] Similarly, to offset or pre-compensate for the effects of other acoustic characteristics of the listening area, the acoustics of the listening area can be detected and analyzed to determine equalization settings for the playback entities to compensate for theacoustics of the listening area (e g., via the spectral analysis and calibration methods disclosed herein). And playback entities within the listening area can be configured with equalization settings that offset or pre-compensate for the acoustic characteristics of the listening area.

[0046] For example, if the listening area tends to undesirably amplify certain acoustic frequencies more than others, the playback entities within the listening area can be configured with equalization settings that attenuate those acoustic frequencies during audio playback to offset the amplification of those frequencies by the listening area. Or if the listening area tends to undesirably attenuate certain acoustic frequencies more than others, the playback entities within the listening area can be configured with equalization settings that amplify those acoustic frequencies during audio playback to offset the attenuation of those frequencies by the listening area.

[0047] Finally, to address coherent addition of low frequency content within the listening area, the extent to which bass frequencies tend to sum within the listening area can be detected, and equalization settings for the playback entities to counteract the bass frequency accumulation can be determined (e.g., via the low frequency analysis and correction methods disclosed herein).

[0048] Thus, some embodiments include, for one or more (or all) playback entities within a playback group (including Area Zone configurations) deployed within a listening area, (1) determining one or more boundary conditions for the playback entity to identify whether the speakers associated with the playback entity are placed near an acoustically reflective surface such as a ceiling, wall, or corner, to determine playback settings (e g., volume and / or equalization settings) to ameliorate the effect of acoustic reflections, (2) performing a spectral analysis of the listening area to determine equalization settings for audio playback to account for the frequency response of the listening area, and perhaps also performing a spatial analysis of the listening area to determine delay settings for different playback entities within the listening area, and (3) performing a low frequency analysis and correction to generate equalization settings to help avoid excessive bass levels within the listening area.A. Boundary Condition Determination and Calibration

[0049] In some embodiments, determining one or more boundary conditions for a playback entity include causing the playback entity to determine its self-response. A playback entity determining its self-response may include, for example, (i) emitting a calibration signal viaone or more speakers associated with the playback entity, and (ii) receiving one or more acoustic reflections of the calibration signal. The acoustic reflections may be obtained via one or more microphones integrated within the playback entities, a microphone of a controller device, and / or any other suitable microphone separate from the playback entities and the controller device that is suitably positioned to measure reflections sufficient for determining a self-response for the playback entity.

[0050] One or more parameters of the measured acoustic reflections can be compared to one or more parameters of an expected acoustic reflection to help identify whether the speakers associated with the playback entity are near an acoustically reflective surface, and in some instances, perhaps how far away the speakers associated with the playback entity are from the acoustically reflective surface. Measuring the acoustic reflections can also identify the acoustic frequencies of the reflections.

[0051] In other examples, the playback entity may use sensor data alone or combination with the self-response to detect a boundary condition. For instance, early reflections of ultrasound signals emitted from and / or detected by one or more ultrasonic transducers may indicate proximity to an acoustically reflective surface (e.g., a wall, comer, ceiling, furniture, or other reflective surface). In such scenarios, the ultrasonic reflections may be used instead of or in addition to the self-response of the playback entity. Using ultrasonic emitters and detectors and / or perhaps other sensor data may be advantageous in implementations that include larger numbers of playback entities because it can help reduce the overall amount of time required to determine boundary conditions for all of the playback entities.

[0052] In still other examples, a graphical user interface may display the location(s) of the speakers associated with the playback entities within the Area Zone configuration so that an installer (or other system operator) can select a subset of playback entities having speakers positioned near reflective surfaces to perform the boundary condition detection and calibration. For example, for an Area Zone having a square grid of speakers located in the ceiling, the installer (or other system operator) can select the playback entities having speakers in each comer of the room to perform the boundary condition detection and calibration while perhaps foregoing boundary condition calibration for playback entities having speakers positioned in the middle of the ceiling (i.e., not near walls or comers). Performing the boundary conditioncalibration for only a subset of the playback entities can, in some instances, reduce the total time to perform the overall calibration procedures.

[0053] After detecting a boundary condition for a playback entity, the playback entity is configured with equalization settings that reduce the gain applied to (or perhaps apply negative to) acoustic frequencies that are reflected by the acoustically reflective surfaces so as to reduce the extent to which the audio heard by listeners is affected by the reflections.B. Spectral and / or Spatial Analysis and Calibration

[0054] In some embodiments, performing the spectral analysis of the listening area to determine equalization settings for audio playback to account for the frequency response of the listening area, and perhaps also performing the spatial analysis of the listening area to determine delay settings for different playback entities within the listening area is the same as or similar to the procedures disclosed and described in one or both of (i) U.S. App. 15 / 211,822, titled, “Spatial Audio Correction,” referred to as Docket No. 16-0402, fded on Jul. 15, 2016, and issued on Oct. 17, 2017, as U.S. Pat. 9,794,710 and / or (ii) U.S. App. 16 / 115,524, titled “Playback Device Calibration,” referred to as Docket No. 18-0401, fded on Aug. 28, 2018, and issued on May 21, 2019, as U.S. Pat. 10,299,061. The entire contents of U.S. Apps. 15 / 211,822 and 16 / 115,524 are incorporated herein by reference.

[0055] For example, conducting the spectral and / or spatial analysis of the listening area in some embodiments includes playing a calibration signal via one or more speakers associated with the playback entities of the Area Zone, and using one or more microphones to capture a recording of the playback of the calibration signal within the listening area. In some embodiments the calibration signal used for the spectral analysis may include an audio signal that varies in frequency over time, e.g., in a frequency sweeping pattern. Sweeping through frequencies over time enables the microphones recording to the playback of the calibration signal to capture the acoustic response of the room as a function of frequency.

[0056] In operation, the recording of the playback of the calibration signal may be obtained via one or more microphones integrated within the playback entities, a microphone of a controller device, and / or any other suitable microphone separate from the playback entities and the controller device that is suitably positioned within the listening area to obtain the recording to the playback of the calibration signal.

[0057] However, in some embodiments, performing the spectral and / or spatial analysis of the listening area may deviate from the procedures disclosed and described in U.S. Apps. 15 / 211,822 and 16 / 115,524 in several ways.

[0058] For example, prior solutions typical include performing a spectral analysis of a room in which one or more playback devices are deployed, and then determining a set of equalization settings for the playback devices in the room. In operation, the spectral analysis determines the acoustic characteristics of the room as whole. But because Area Zones may include greater numbers of playback entities and be deployed in larger spaces, some embodiments include performing several localized spectral analyses of different areas within a room, and then customizing equalizations for the different playback entities within the different areas of the room perhaps on an area-by-area basis. In this manner, such embodiments include performing multiple localized spectral analyses for a single room rather than prior approaches which typically include performing a single global spectral analysis for the room as a whole.

[0059] Additionally, in some instances, rather than utilizing a microphone incorporated within a playback device or a controller device, some embodiments may instead utilize a microphone that is separate from any particular playback device or controller device, such as a microphone having a larger and / or more sensitive transducer that can obtain better quality audio recordings than microphones typically implemented within playback devices and / or controller devices.

[0060] Further, in some instances, rather than conducting the spatial analysis to determine playback timing delays to direct the time-of-arrival of audio to one spatial location within the listening area, some embodiments may instead conduct the spatial analysis to determine playback timing delays direct the time-of-arrival of audio to several different spatial locations within the listening area.C. Low Frequency Analysis and Correction

[0061] In some embodiments, performing a low frequency analysis and correction to generate equalization settings to help avoid excessive bass levels within the listening area includes all of the playback entities simultaneously (or at substantially the same time) emitting a sound signal from their speakers. The sound signal may comprise a broadband noise signal, e.g.,pink noise, white noise, or other suitable broadband noise signal. In some instances, the sound signal may comprise one or more bursts of broadband noise.

[0062] Based on the noise emitted from the speakers of the playback entities, a decay time (e.g., RT60, which is defined as the measure of the time after the sound source ceases that it takes for the sound pressure level to reduce by 60 dB) is determined to assess an overall bass level output within the listening area. The audio for determining the decay may be obtained via one or more microphones integrated within the playback entities, a microphone of a controller device, and / or any other suitable microphone separate from the playback entities and the controller device.

[0063] After measuring the decay time, one or more equalization settings (e.g., a bass level) can be adjusted on one or more playback entities within the listening area. In scenarios where the decay time is measured at different locations within the listening area, equalization settings can be tailored to each of the different locations based on the measured decay time for different locations. In operation, tailoring the equalization settings to the different locations may include using different bass level settings for the different playback entities at the different locations within the listening area so as to avoid excessive bass build up within the listening area in general, and within different locations within the listening area in particular.D. Overview of Calibration Procedures

[0064] The three features identified above may be performed in different orders. For example, instead of (1) determining boundary conditions and calibrations, (2) performing the spectral (and perhaps spatial) analysis and calibration of the listening area, and (3) performing the low frequency analysis and correction as described above, some embodiments may instead (1) determine boundary conditions, (2) perform the low frequency analysis and correction, and (3) perform the spectral (and perhaps spatial) analysis of the listening area. Still other embodiments may (1) perform the low frequency analysis and correction, (2) determine boundary conditions, and (3) perform the spectral (and perhaps spatial) analysis of the listening area. Other orderings are possible, too.

[0065] Additionally, some embodiments include performing only one of the three processes, but not necessarily the other two. Likewise, some embodiments include performingonly two of the three processes. And of course, some embodiments include performing all three processes.

[0066] For example, some embodiments include a computing device that is configured to, for an area zone configuration comprising a first playback entity and one or more other playback entities located in a listening environment: (i) cause the first playback entity to emit a first audio signal via one or more speakers associated with the first playback entity; (ii) obtain a first recording from at least one microphone associated with the first playback entity, wherein the first recording comprises reflections of the first audio signal from one or more objects in the listening environment; (iii) determine a self-response of the first playback entity based on the first recording; (iv) cause the first playback entity to emit a second audio signal via the one or more speakers while each of the one or more other playback entities in the area zone configuration also emit the second audio signal, wherein the second audio signal is different than the first audio signal; (v) obtain a second recording of the second audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration; (vi) determine a bass level in the listening environment based on the second recording; and (vii) when the determined bass level in the listening environment differs from a target bass level for the listening environment by more than a threshold amount, adjust a bass equalization level of at least one playback entity in the area zone configuration.

[0067] The computing device may be any of (i) a computing device separate from the first playback entity and the one or more other playback entities in the area zone configuration, (ii) the first playback entity, or (iii) one of the one or more other playback entities in the area zone configuration.

[0068] In some embodiments, the computing device is further configured to, for the area zone configuration comprising the first playback entity and the one or more other playback entities located in the listening environment: (i) cause the first playback entity to emit a third audio signal via the one or more speakers while each of the one or more other playback entities in the area zone configuration also emit the third audio signal, wherein the third audio signal is different than the first audio signal and the second audio signal; (ii) obtain a third recording of the third audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration; (iii) determine an acoustic response of the listening environment based on the third recording; and (iv) when the determined acousticresponse of the listening environment differs from a target acoustic response for the listening environment by more than a threshold amount, adjust one or more equalization levels of at least one playback entity in the area zone configuration.

[0069] Some embodiments include a system comprising a first playback entity and one or more additional playback entities configured to play audio in a listening environment. In operation, the first playback entity is configured to: (i) emit a first audio signal via one or more speakers associated with the first playback entity; (ii) obtain a first recording from at least one microphone associated with the first playback entity, wherein the first recording comprises reflections of the first audio signal from one or more objects in the listening environment; (iii) determine one or more boundary conditions of the first playback entity based on the first recording comprising the reflections of the first audio signal; (iv) emit a second audio signal via the one or more speakers associated with the first playback entity while each of the one or more other playback entities in the system also emit the second audio signal, wherein the second audio signal is different than the first audio signal; (v) obtain a second recording of the second audio signal emitted by the first playback entity and each of the one or more other playback entities in the system; (vi) determine a bass level in the listening environment based on the second recording; and (vii) when the determined bass level in the listening environment differs from a target bass level for the listening environment by more than a threshold amount, adjust a bass equalization level of at least the first playback entity.

[0070] In some system embodiments, the one or more other playback entities in the system comprise a second playback entity. In operation the second playback entity is configured to perform some of the same functions as the first playback entity. For example, the second playback entity is configured to: (i) emit a self-response audio signal via one or more speakers associated with the second playback entity; (ii) obtain a reflection recording from at least one microphone associated with the second playback entity, wherein the reflection recording comprises reflections of the self-response audio signal from one or more objects in the listening environment; (iii) determine a self-response of the second playback entity based on the reflection recording from at least one microphone associated with the second playback entity; (iv) determine whether one or more aspects of the reflections in the reflection recording for the second playback entity differ from one or more expected aspects of the reflections for the second playback entity by more than a threshold amount; and (v) when one or more aspects of thereflections in the reflection recording for the second playback entity differ from one or more expected aspects of the reflections by more than a threshold amount, adjust one or more equalization parameters of at least the second playback entity.

[0071] Additional aspects of the disclosed systems and methods and disclosed and described in further detail herein.E. Relevant Nomenclature

[0072] Given the new concepts described herein, certain terminology is introduced and used for explaining various example features and embodiments. However, it should be understood that such terminology and the use thereof may be uniquely applicable in at least some respects to the examples described herein, such as when describing both existing concepts and new concepts.

[0073] For example, to help illustrate aspects of some embodiments, a playback device as used herein sometimes refers to a single, physical hardware device that is configured to play audio. Such a playback device includes one or more network interfaces, one or more processors, and tangible, non-transitory computer-readable media storing program instructions that are executed by the one or more processors to cause the playback device to perform the playback device features and functions described herein.

[0074] In some embodiments, a playback device includes integrated speakers. In other embodiments, a playback device includes speaker outputs that connect to external speakers. In still further embodiments, a playback device includes a combination of integrated speakers and speaker outputs that connect to external speakers.

[0075] In some instances, a playback device may include a set of headphones. In some instances, a playback device may include a smartphone, tablet computer, laptop / desktop computer, smart television, or other type of device configurable to play audio content.

[0076] In some embodiments, a playback device comprises one or more microphones configured to receive voice commands. In some instances, playback devices with microphones are referred to herein as Networked Microphone Devices (NMDs). In some NMD embodiments, the NMD is configured to perform any (or all) of the playback device functions disclosed herein.

[0077] Additional details about playback devices consistent with some example embodiments are disclosed and described herein.

[0078] As another example, a logical player as used herein sometimes refers to a logical playback entity implemented by one or more physical playback devices to act as a single entity to play one stream of audio. In some instances where the one stream of audio comprises multichannel audio having two or more channels and the player includes two or more channel outputs, playing the multichannel audio includes the player playing the two or more channels via two or more corresponding channel outputs. In some instances, a single physical playback device implements a single logical player. In other instances, several physical playback devices may be configured to implement a single logical player. In still further instances, one physical playback device (i.e., a Multi-Player Playback Device) may implement several logical players.

[0079] Additional details about logical players consistent with some example embodiments are disclosed and described herein

[0080] As another example, a Multi-Player Playback Device as used herein sometimes refers to a type of physical playback device comprising multiple configurable audio outputs, one or more processors, and tangible, non-transitory computer readable media storing program instructions that are executed by the one or more processors to cause the Multi-Player Playback Device to perform the Multi-Player Playback Device features and functions described herein. In scenarios where a playback device might be referred to as a zone player, the Multi-Player Playback Device may sometimes be referred to as a Multi-Zone Player.

[0081] In operation, the multiple audio outputs can be grouped together in different combinations to implement one or more logical players. In some embodiments, a single Multi- Player Playback Device is configurable to implement from one to eight logical players. Multi- Player Playback Devices according to some embodiments include eight configurable audio outputs. Multi-Player Playback Devices according to other embodiments include fewer than eight configurable audio outputs or more than eight configurable audio outputs. For example, in some embodiments, a Multi-Player Playback Device may include anywhere from two to six, eight, twelve, sixteen, eighteen, twenty four, or more configurable audio outputs. Multi-Player Playback Devices with more or fewer audio outputs than those specifically identified herein are possible as well.

[0082] Additional details about Multi-Player Playback Devices consistent with some example embodiments are disclosed and described herein and in U.S. Provisional App.63 / 502,347, titled “Area Zones,” referred to as Docket No. 22-1002p, fded on May 15, 2023, the contents of which are incorporated herein by reference.

[0083] As another example, a playback entity as used herein sometimes refers to a logical or physical entity configured to play audio. Playback entities include physical playback devices, physical Multi-Player Playback Devices, and logical players that are implemented via one or more Multi-Player Playback Devices.

[0084] Additional details about playback entities consistent with some example embodiments are disclosed and described herein and in U.S. Provisional App. 63 / 502,347, titled “Area Zones,” referred to as Docket No. 22-1002p, filed on May 15, 2023, the contents of which are incorporated herein by reference.

[0085] As another example, a zone (sometimes referred to herein as a playback zone or a bonded zone) as used herein sometimes refers to a logical container of one or more physical playback devices that are managed together. But in some instances, and some of the examples described herein, a zone may include only a single, physical playback device. Thus, in operation, a zone may include any one or more playback devices managed as a logical zone entity, including, for example, (i) a single playback device managed as a logical zone entity, (ii) a group of playback devices managed as a logical zone entity, including but not limited to any of(a) a bonded zone that includes two or more playback devices configured to play the same audio,(b) a bonded pair of two playback devices configured to play the same audio, (c) a stereo pair of two playback devices where one of the playback devices is configured to play a left channel of stereo audio content and the other playback device is configured to play a right channel of stereo audio content, or (d) a home theater zone that includes two or more playback devices configured to play home theater and / or surround sound audio content.

[0086] In this manner, in some examples, a zone is a type of logical entity implemented by one or more playback devices. When the zone includes two or more playback devices, the two or more playback devices play one stream of audio content. In some instances, the one stream of audio content played by the zone comprises multichannel audio. In some zone scenarios that include two or more playback devices where the one stream of audio content comprises multichannel audio, each playback device (of the two or more physical playback devices) may be configured to play a different channel of the multichannel audio content.

[0087] For example, a first physical playback device in the zone may be configured to play a left channel of the audio content and a second physical playback device in the zone may be configured to play a right channel of the audio content. This type of example zone configuration is sometimes referred to as a stereo pair.

[0088] In another example, a first playback device in the zone may play a left channel, a second playback device in the zone may play a right channel, and a third playback device in the zone may play a subwoofer channel. This type of example zone configuration is sometimes referred to as a home theater zone.

[0089] In some existing zone configurations, each of the individual playback devices within the zone communicate with each other in a fully-connected control plane configuration to exchange commands, configuration information, events, and state information to each other via dedicated websocket connections between each pair of playback devices within the zone.

[0090] Additional details about zones consistent with some example embodiments are disclosed and described herein and in U.S. Provisional App. 63 / 502,347, titled “Area Zones,” referred to as Docket No. 22-1002p, filed on May 15, 2023, the contents of which are incorporated herein by reference.

[0091] A playback group (sometimes referred to herein simply as a group) is a logical container of two or more logical or physical playback entities. Logical and physical entities that can grouped into a playback group include: (i) a zone, (ii) a playback device, (iii) a Multi-Player Playback Device, and / or (iii) a logical player.

[0092] A playback group differs from a zone in a few ways. First, a zone includes one or more playback devices, whereas a playback group includes two or more logical or physical playback entities (i.e., zones, playback devices, Multi-Player Playback Device, or players).Second, when one or more playback devices are configured into a zone, the playback system treats the zone as a single logical entity even though the zone may include several physical playback devices. By contrast, when two or more playback devices are configured into a group, the playback system manages each playback device separately even though each of the playback entities within the playback group are playing the same audio stream.

[0093] Similar to playback devices configured into a zone, all of the playback entities within a playback group are configured to play the same audio stream. In some instances, the audio stream comprises multichannel audio. In some scenarios where the audio streamcomprises multichannel audio, each zone, playback device, Multi-Player Playback Device, or player in the playback group is configured to play the same set of audio channels. In some example playback group implementations, the playback group includes a group coordinator and one or more group members. The group coordinator sources audio for the playback group and manages certain configuration and control functions for the group member(s) in the playback group. However, other playback group implementations may distribute audio among the group and handle configuration and control functions differently.

[0094] Similar to a zone configuration (described above), and in contrast to an Area Zone configuration (described further herein), each of the individual playback devices within a playback group communicate with each other in a fully-connected configuration to exchange commands, configuration information, events, and state information with each other via dedicated websocket connections (or similar communication links / sessions) between each of the playback devices within the playback group. For example, a playback group with four playback devices in some playback group configurations would include six dedicated websocket connections (or similar communication links / sessions) for a fully-connected mesh between all four playback devices, e.g., dedicated connections between (i) playback device 1 and playback 2, (ii) playback device 1 and playback device 3, (iii) playback device 1 and playback device 4, (iv) playback device 2 and playback device 3, (v) playback device 2 and playback device 4, and (v) playback device 3 and playback device 4.

[0095] Additional details about playback groups consistent with some example embodiments are disclosed and described herein and in U.S. Provisional App. 63 / 502,347, titled “Area Zones,” referred to as Docket No. 22-1002p, filed on May 15, 2023, the contents of which are incorporated herein by reference.A. Area Zone

[0096] An Area Zone is a new type of playback grouping. An Area Zone includes a set of two or more logical or physical playback entities grouped together. The playback entities that can be grouped together into an Area Zone include: (i) a zone, (ii) a playback device, (iii) a Multi-Player Playback Device, (iv) a logical player, and / or (iv) a playback group.

[0097] Similar to some zone and a playback group configurations (described above), all of the playback entities within an Area Zone are configured to play one audio stream. In someinstances, the one audio stream comprises a multichannel audio stream. In some embodiments described herein, an Area Zone includes an Area Zone Primary and one or more Area Zone Secondaries, each of which is described further herein.

[0098] In some embodiments, and similar to some zone configurations (described above), the playback system manages all of the playback entities within an Area Zone as a single logical playback entity. And similar to some playback group configurations (described above), in some embodiments, the individual playback entities can be managed and configured independently of each other.

[0099] One difference between some zone implementations (described above) on the one hand, and an Area Zone on the other, is that in some existing systems, prior zone configurations cannot be saved, and then activated or deactivated during operation of the playback system.With some prior zone implementations, the zone is configured typically when the playback devices forming the zone are added to the playback system. In contrast to those prior zone implementations, with some Area Zone embodiments, an individual playback entity can save several different Area Zone configurations and switch between operating in each of the different saved Area Zone configurations.

[0100] Another difference between zones and playback groups (described above) on the one hand, and an Area Zone on the other, is that unlike prior zone and playback group configurations, the playback entities within an Area Zone do not all communicate with each other in a fully-connected control plane configuration to exchange commands, configuration information, events, and state information with each other. In some examples, this fully- connected control plane is implemented via dedicated websocket connections between each pair of playback entities within the Area Zone.

[0101] Instead, the Area Zone Primary communicates with each Area Zone Secondary (and each Area Zone Secondary communicates with the Area Zone Primary) to exchange commands, configuration information, events, and state information in a hierarchical control plane. In contrast to how group members within a playback group maintain a fully-connected mesh between each other (e.g., via dedicated websocket connections in some instances) to facilitate communication between the different group members, in normal operation, the Area Zone Secondaries within the same Area Zone typically do not communicate with each other, and Area Zone Secondaries typically do not communicate with any other playback entity in aplayback system other than their corresponding Area Zone Primary. Instead, the Area Zone Secondaries communicate with the Area Zone Primary which can, in turn, facilitate any exchange of commands, configuration information, events, or state information that may need to occur between two Area Zone Secondaries.

[0102] Additional details about Area Zones consistent with some example embodiments are disclosed and described herein and in U.S. Provisional App. 63 / 502,347, titled “Area Zones,” referred to as Docket No. 22-1002p, filed on May 15, 2023, the contents of which are incorporated herein by reference.B. Area Zone Primary

[0103] An Area Zone Primary is a player (e.g., a playback device or Multi-Player Playback Device) that is configured to handle audio sourcing and control signaling on behalf of itself and all of the Area Zone Secondaries within an Area Zone. In some embodiments, the Area Zone Primary may also perform one or more (or all) functions of a Global State Aggregator as described in U.S. Provisional App. 63 / 377,978, titled “Global State Service,” referred to as Docket No. 23-0306p, filed on Sep. 30, 2022.

[0104] For media, in some implementations, the Area Zone Primary is configured to function as an audio sourcing device for itself and all of the Area Zone Secondaries within the Area Zone.

[0105] For control signaling, in some implementations, the Area Zone Primary communicates with each Area Zone Secondary to exchange commands, configuration information, events, and state information to implement playback, configuration, and control functions (e.g., volume, mute, playback start / stop, queue management, configuration management and updates) for the Area Zone.

[0106] In addition to exchanging commands, configuration information, events, and state information with each Area Zone Secondary within the Area Zone, each Area Zone Primary in some implementations is also configured to exchange commands, configuration information, events, and state information with (i) each (and every) other Area Zone Primary in the playback system and (ii) any controller device(s) or system(s) configured for controlling operation of the playback system. However, rather than exchanging commands, configuration information, events, and state information with each (and every) other Area Zone Primary in the playbacksystem, in some embodiments each Area Zone Primary in some implementations is configured to exchange commands, configuration information, events, and state information with one or more Brokers (not every other Area Zone Primary) in the playback system. The use of Brokers in this an other manners is described in more detail in U.S. Provisional App. 63 / 377,978, titled “Global State Service,” referred to as Docket No. 23-0306p, filed on Sep. 30, 2022.

[0107] One example scenario that illustrates how an Area Zone Primary exchanges commands, configuration information, events, and state information with Area Zone Secondaries and / or controller device(s) and / or controller system(s) is where, after a playback entity configured as the Area Zone Primary receives a request for configuration or operational information about one or more playback entities within the Area Zone from a requesting device (e.g., a controller device, a controller system, or perhaps another playback entity in the playback system), the playback entity configured as the Area Zone Primary provides the requested information to the requesting device on behalf of the Area Zone.

[0108] For example, in response to a request for a listing of playback entities in the Area Zone received from a requesting device, the Area Zone Primary transmits a listing of playback entities within the Area Zone to the requesting device. In another example, in response to a request for configuration information about a particular Area Zone Secondary received from a requesting device, and to the extent that the Area Zone Primary does not already have the requested configuration information for the particular Area Zone Secondary, the Area Zone Primary obtains the requested configuration information from the Area Zone Secondary. And regardless of whether the Area Zone Primary already had the configuration information for the particular Area Zone Secondary or acquired the requested configuration information from the particular Area Zone Secondary, the Area Zone Primary provides the requested configuration information for the particular areas on secondary to the requesting device.

[0109] Additional details about Area Zone primaries consistent with some example embodiments are disclosed and described herein and in U.S. Provisional App. 63 / 502,347, titled “Area Zones,” referred to as Docket No. 22-1002p, filed on May 15, 2023, the contents of which are incorporated herein by reference.C. Area Zone Secondary

[0110] An Area Zone Secondary is a logical or physical entity in the Area Zone that is not the Area Zone Primary for the Area Zone. In operation, each Area Zone Secondary within an Area Zone is configured to communicate with the Area Zone Primary to exchange commands, configuration information, events, and state information to implement playback, configuration, and control functions (e.g., volume, mute, playback start / stop, queue management, configuration management and updates) for the Area Zone.[0U1] As mentioned earlier, in normal operation, an Area Zone Secondary typically does not communicate with any other Area Zone Secondary within the Area Zone or any other playback entity within a playback system. However, an Area Zone Secondary in some instances may receive commands and / or inquiries from a controller device, and in some instances can establish a communication session with another Area Zone Secondary to exchange data in some circumstances described herein.

[0112] One example scenario that illustrates how an Area Zone Secondary exchanges commands, configuration information, events, and state information with its Area Zone Primary is where, after a playback entity configured as an Area Zone Secondary receives a request for configuration or operational information about one or more playback entities within the Area Zone from a requesting device (e.g., a controller device, a controller system, or perhaps another playback entity in the playback system), the playback entity configured as the Area Zone Secondary forwards the received request to the Area Zone Primary. In some instances, the Area Zone Primary responds to the requesting device to provide the requested information.

[0113] For example, in response to a request for a listing of playback entities in the Area Zone received from a requesting device, the Area Zone Secondary forwards the request to the Area Zone Primary, and the Area Zone Primary transmits a listing of playback entities within the Area Zone to the requesting device. In another example, in response to a request for configuration information about a particular Area Zone Secondary received from a requesting device (including a request about itself), the Area Zone Secondary forwards the request to the Area Zone Primary, and the Area Zone Primary provides the requested information to the requesting device.

[0114] Additional details about Area Zone secondaries consistent with some example embodiments are disclosed and described herein and in U.S. Provisional App. 63 / 502,347, titled“Area Zones,” referred to as Docket No. 22-1002p, filed on May 15, 2023, the contents of which are incorporated herein by reference.D. Controller Devices

[0115] A controller device (sometimes referred to herein as a controller or a control device) is a computing device or computing system with one or more processors and tangible, non-transitory computer readable media storing program instructions executable by the one or more processors to execute the controller device features and functions described herein. In some scenarios, a controller device is a smartphone, tablet computer, laptop computer, desktop computer, smartwatch, or similar computing device configured to execute a software user interface for configuring and controlling playback entities within a playback system. In some scenarios, a controller device may include one or more cloud server systems configured to communicate with one or more playback entities to configure and control the playback system.

[0116] In operation, user inputs associated with commands for configuring and controlling playback entities within a playback system can take a variety of forms, including but not limited to (i) physical inputs (e.g., actuating physical controls like knobs, sliders, buttons and so forth) (ii) software user interface inputs (e.g., inputs on a touch screen or similar graphical user interface), (iii) voice inputs, (iv) inputs received from another playback entity in the playback system (e.g., in the form of signaling from the playback entity to the controller in connection with effectuating configuration and control commands) and / or (v) any other type of input in any other form now known or later developed that is sufficient for conveying commands for configuring and controlling playback entities.

[0117] Additional details about area controller devices consistent with some example embodiments are disclosed and described herein and in U.S. Provisional App. 63 / 502,347, titled “Area Zones,” referred to as Docket No. 22-1002p, filed on May 15, 2023, the contents of which are incorporated herein by reference.E. Hierarchical Control Plane for Command and Control Signaling

[0118] As mentioned above, one aspect of the Area Zone embodiments disclosed herein is a hierarchical control plane.

[0119] Exchanging command and control information via control plane implementations disclosed herein differs from prior command and control distribution schemes such as the onesdisclosed in U.S. App. 13 / 489,674, titled “Device Playback Failure Recovery and Redistribution,” filed on Jun. 6, 2012, and issued on Dec. 2, 2014, as U.S. Pat. 8,903,526, the entire contents of which are incorporated herein by reference. The ‘674 application describes a command and control distribution scheme where a confirmed communication list (CCL) is be generated to facilitate communication between playback devices within a playback system. In one example, the CCL is a list of all playback devices in a zone configuration, where the CCL is ordered according to an optimal routing using the least number of hops or transition points through the network between the playback devices. In another case, the CCL is generated without consideration of network routing metrics. In either case, command and control data is passed from playback device to playback device within the zone configuration following the order in the CCL in a linear or serial manner. In one example, a first playback device sends a command to a second playback device in the CCL, and the second playback device in the CCL sends the command to a third playback device in the CCL, and so on until the command reaches its destination (i.e., a playback device in the CCL). For commands to be processed by all playback devices in the zone configuration, the commands are routed from playback device to playback device in the order specified in the CCL until every playback device in the CCL has received the commands. This arrangement is simple to execute, provides reliable transmission of information from playback device to playback device, and tends to work quite well for zones with a few playback devices.

[0120] However, for Area Zones with a lot of playback entities, the CCL-based approach can become impractical because it can take too long to distribute commands to a large number of playback entities or even to send a command to a single playback entity since every command is routed through the group on a serial, hop-by-hop basis according the CCL.

[0121] In contrast to the above-described CCL approach and flat control plane implementations where many (or all) playback entities within a playback system communicate directly with each other to exchange configuration and control information throughout the playback system, some Area Zone embodiments disclosed herein employ a hierarchical control plane where commands, configuration information, events, and state information are exchanged only between the Area Zone Primary and each Area Zone Secondary within the Area Zone. For playback systems that may include several Area Zones, the Area Zone Primaries may exchange commands, configuration information, events, and state information with each other.

[0122] In the Area Zone embodiments disclosed herein, and in contrast to flat, fully- meshed control plane implementations, the Area Zone Secondaries typically do not exchange commands, configuration information, events, and state information to each other via dedicated websocket connections (or similar communications links or sessions) with each other, except perhaps in a few rare instances described herein.

[0123] In some scenarios, any commands, configuration information, events, and state information to be sent from a first Area Zone Secondary to a second Area Zone Secondary are sent from the first Area Zone Secondary to the Area Zone Primary. The Area Zone Primary then, in turn, (i) processes the command(s), configuration information, event(s), and / or state information received from the first Area Zone Secondary and instructs and / or updates the second Area Zone Secondary accordingly, or (ii) forwards the command(s), configuration information, event(s), and / or state information to the second Area Zone Secondary, as necessary.

[0124] Some Area Zone embodiments additionally or alternatively employ a Configuration and Command (C&C) group (e.g., a multicast group) for distributing commands, configuration information, events, and state information to playback entities within the Area Zone.

[0125] In some instances, the Area Zone Primary creates the C&C group for the Area Zone (or joins the C&C group as a publisher), and provides information required to join and / or subscribe to the C&C group to each Area Zone Secondary. The Area Zone Secondaries join and / or subscribe to the C&C group to receive control information. In such embodiments, the Area Zone Primary publishes commands, configuration information, events, and state information to the C&C group, and each Area Zone Secondary subscribed to the C&C group receives the commands, configuration information, events, and state information that the Area Zone Primary publishes to the C&C group. In some embodiments, Area Zone Secondaries may also publish certain command, configuration, event, and state information to the Area Zone C&C group. In operation, each when an Area Zone Secondary receives a command via the C&C group that requires some action on behalf of the Area Zone Secondary, the Area Zone Secondary executes the received command.

[0126] Additional details about hierarchical control plane features and functionality for Area Zone configurations consistent with some example embodiments are disclosed and described herein and in U.S. Provisional App. 63 / 502,347, titled “Area Zones,” referred to asDocket No. 22-1002p, filed on May 15, 2023, the contents of which are incorporated herein by reference.F. Media and Timing Distribution

[0127] Another aspect of some Area Zone embodiments disclosed herein is how audio content and playback timing are distributed to individual playback entities within the Area Zone. In some existing zone or playback group configurations, the playback device designated as the audio sourcing device is configured to transmit audio content and playback timing for the audio content to each other playback device in the zone or playback group via unicast transmissions from the audio sourcing device to each other playback device in the zone or playback group.

[0128] However, in some Area Zone embodiments disclosed herein, the Area Zone Primary is configured to transmit audio content and playback timing for the audio content to a media distribution group (e.g., a multicast group), and Area Zone Secondaries subscribe to the media group to receive the audio content and playback timing, e.g., each Area Zone Secondary joins the media multicast group and receives audio content and playback timing via the media multicast group.

[0129] Some embodiments employ a hybrid unicast / multicast approach where the Area Zone Primary is configured to (i) distribute audio content and playback timing via unicast transmissions to Area Zone Secondaries that are wirelessly connected to the playback system, e.g., via WiFi, Bluetooth, or other suitable wireless connection, and (ii) distribute audio content and playback timing via multicast transmissions to Area Zone Secondaries that are connected to the playback system via wired connections, e.g., Ethernet, Power over Ethernet (PoE), Universal Serial Bus (USB), or other suitable wired connection.

[0130] Some Area Zone embodiments employ a Media and Timing (M&T) group (e.g., a multicast group) for distributing audio content and playback timing to playback entities within the Area Zone. In some Area Zone embodiments, clock timing is also distributed to the playback entities within the Area Zone via the M&T group. In some instances, the Area Zone Primary creates the M&T group for the Area Zone (or joins the M&T group as a publisher), and provides information required to join and / or subscribe to the M&T group to each Area Zone Secondary. The Area Zone Secondaries join and / or subscribe to the M&T group to receive audio content, playback timing, and in some instances, clock timing. In such embodiments, the Area ZonePrimary publishes the audio content, playback timing, and clock timing to the M&T group, and each Area Zone Secondary subscribed to the M&T group receives the audio content, playback timing, and clock timing that the Area Zone Primary publishes to the M&T group.

[0131] Additional details about media and timing data for Area Zone configurations consistent with some example embodiments are disclosed and described herein and in U.S. Provisional App. 63 / 502,347, titled “Area Zones,” referred to as Docket No. 22-1002p, filed on May 15, 2023, the contents of which are incorporated herein by reference.G. Storing and Recalling Area Zone Configuration Information

[0132] Another aspect of some Area Zone embodiments disclosed herein includes the Area Zone configuration data and how the Area Zone configuration data is stored and recalled to activate an Area Zone configuration.

[0133] In contrast to some prior playback group configurations and zone configurations where every playback device maintains an up-to-date version of the group / zone configuration data, Area Zone configurations according to some embodiments include storing the Area Zone configuration information in a single Area Zone configuration package. In some instances, the Area Zone configuration package includes separate Area Zone configuration files for each playback entity in the Area Zone. In some instances, the Area Zone configuration package is stored at one or more of (i) the Area Zone Primary, (ii) a controller device, and / or (iii) a cloud server system.

[0134] The Area Zone configuration information in some embodiments includes information about the Area Zone configuration, including but not limited to one or more (or all) of (i) a name that identifies the Area Zone, (ii) the name of the playback entity that is to function as the Area Zone Primary once the Area Zone is activated, (iii) the name of each playback entity in the Area Zone (i.e., the name of each playback device, Multi-Player Playback Device, player, zone, and playback group in the Area Zone, as applicable), (iv) playback calibration settings for each playback entity in the Area Zone, including but not limited to playback calibration settings determined according to the methods and processes disclosed and described herein, (v) for Area Zones configured to play multichannel audio, a channel map that defines the channel or channels that each playback entity is configured to play once the Area Zone has been activated, including at least in some instances, which channel each output port of each playback entity is configuredto play while the Area Zone is active, and (vi) for individual playback entities, one or more configuration setting(s) that define certain Area Zone specific behavior of the playback entity while the Area Zone is active.

[0135] Examples of Area Zone specific behavior while the Area Zone is active include: (i) behavior of the playback entity in response to receiving a volume control command, a playback control command (e.g., play / pause / skip / etc.), and / or other user command via a physical interface on (or associated with) the playback entity, and (ii) behavior of the playback entity in response to receiving a voice command via an associated microphone.

[0136] Additional details about storing and recalling Area Zone configuration consistent with some example embodiments information are disclosed and described herein and in U.S. Provisional App. 63 / 502,347, titled “Area Zones,” referred to as Docket No. 22-1002p, filed on May 15, 2023, the contents of which are incorporated herein by referenceH. Related Art

[0137] The above-described example configurations as well as additional and alternative example configurations are described in more detail herein. Aspects of the features and functions implemented in the example configurations disclosed herein differ from known implementations in several ways.

[0138] For example, with respect to streaming audio content and playback timing to individual playback devices, U.S. App. 10 / 816,217, titled “System And Method For Synchronizing Operations Among A Plurality Of Independently Clocked Digital Data Processing Devices,” referred to as Docket No. 04-0401, filed on Apr. 1, 2004, and issued on Jul. 31, 2012, as U.S. Pat. 8,234,395, describes, inter alia, independently-clocked playback devices that are configured to play audio content in synchrony with each other based on playback timing and clock timing information. However, U.S. App. 10 / 816,217 does not describe the Area Zone configurations, Multi-Player Playback Devices, playback entity implementations, and / or the calibration procedures described herein. The entire contents of U.S. App. 10 / 816,217 are incorporated herein by reference.

[0139] Similarly, with respect to groupwise control of groups of playback devices within a playback system, U.S. App. 10 / 861,653, titled “Method And Apparatus For Controlling Multimedia Players In A Multi -Zone System,” referred to as Docket No. 04-0601, filed on Jun.5, 2004, and issued on Aug. 4, 2009, as U.S. Pat. 7,571,014, describes, inter alia, groupwise control of multiple playback devices, including groupwise volume control of playback devices configured within a group of playback devices configured to play audio content in synchrony with each other, sometimes referred to as a synchrony group. However, U.S. App. 10 / 861,653 does not describe does not describe the Area Zone configurations, Multi-Player Playback Devices, playback entity implementations, and / or the calibration procedures described herein. The entire contents of U.S. App. 10 / 861,653 are incorporated herein by reference.

[0140] Other aspects of groupwise control of groups of playback devices are described in U.S. App. 13 / 910,608, titled “Satellite Volume Control,” referred to as Docket No. 13-0413, filed on June. 5, 2013, and issued on Sep. 6, 2016, as U.S. Pat. 9,438,193. For example, U.S. App. 13 / 910,608 discloses, inter alia, controlling playback volume of grouped playback devices, including propagating a volume adjustment received via a first playback device to other playback devices that have been grouped with the first playback device. However, U.S. App. 13 / 910,608 does not describe does not describe the Area Zone configurations, Multi-Player Playback Devices, playback entity implementations, and / or the calibration procedures described herein. The entire contents of U.S. App. 13 / 910,608 are incorporated herein by reference.

[0141] Further, several earlier-filed applications describe aspects of specialized groupings of playback devices within a playback system. For example, U.S. App. 13 / 013,740, titled “Controlling and grouping in a multi-zone media system,” referred to as Docket No. 11- 0101, filed on Jan. 25, 2011, and issued on Dec. 1, 2015, as U.S. Pat. 9,202,509, describes, inter alia, configuring and operating two playback devices in a “paired” configuration such as a “stereo pair” configuration, where one playback device is configured to play a right stereo channel of audio content and the other playback device is configured to play a left stereo channel of audio content.

[0142] Similarly, U.S. App. 13 / 083,499, titled “Multi-Channel Pairing In A Media System,” referred to as Docket No. 11-0401, filed on Apr. 8, 2011, and issued on Jul. 22, 2014, as U.S. Pat. 8,788,080, describes, inter alia, configuring and operating multiple playback devices in a consolidated mode, where two or more playback devices can be grouped into a consolidated playback device which can then be further grouped with one or more other playback devices and / or one or more other consolidated playback devices.

[0143] Further, U.S. App. 13 / 632,731, titled, “Providing A Multi-Channel And A MultiZone Audio Environment,” referred to as Docket No. 12-0802, filed on Oct. 1, 2012, and issued on Dec. 6, 2016, as U.S. Pat. 9,516,440, describes, inter alia, configuring playback devices to play different types of audio content (e.g., home theater audio vs. music) according to different playback timing arrangements (e.g., with low-latency vs. with ordinary latency).

[0144] Additionally, U.S. App. 14 / 731,119, titled, “Dynamic Bonding ofPlayback Devices,” referred to as Docket No. 15-0301, filed on Jun. 4, 2015, and issued on Jan. 9, 2018, as U.S. Pat. 9,864,571, discloses, inter alia, dynamic bonding scenarios and playback devices that are “sharable” among different zones. And U.S. App. 14 / 997,269, titled, “System Limits Based on Known Triggers,” referred to as Docket 15-1104, filed on Jan. 15, 2016, and issued on Feb. 20, 2018, as U.S. Pat. 9,898,245, describes, inter alia, methods of setting up multiple playback devices.

[0145] Further still, U.S. Provisional App. 63 / 377,948, titled “Playback System Architecture,” referred to as Docket No. 21-0703p, filed on Sep. 30, 2022, describes, inter alia, multi-tier hierarchical playback systems comprising several playback devices, including methods and processes of distributing audio and control signaling between and among playback devices with the multi-tier hierarchical playback system.

[0146] However, U.S. Apps. 13 / 013,740; 13 / 083,499; 13 / 632,731; 14 / 731,119;14 / 997,269; and 63 / 377,948 do not describe the Area Zone configurations, Multi-Player Playback Devices, playback entity implementations, and / or the calibration procedures described herein. The entire contents of U.S. Apps. 13 / 013,740; 13 / 083,499; 13 / 632,731; 14 / 731,119;14 / 997,269; and 63 / 377,948 are incorporated herein by reference.

[0147] Additionally, several other advancements over time have improved the overall functionality and usability of playback devices configured in groups for synchronous playback of audio content.

[0148] With respect to managing groups of playback devices configured for groupwise playback of audio content, U.S. App. 14 / 042,001, titled “Coordinator Device for Paired or Consolidated Players,” referred to as Docket No. 13-0812, filed on Dep. 30, 2013, and issued on Mar. 15, 2016, as U.S. Pat. 9,288,596, and U.S. App. 14 / 041,989, titled “Group Coordinator Device Selection,” referred to as Docket No. 13-0815, filed on Sep. 30, 2013, and issued on May 16, 2017, as U.S. Pat. 9,654,545 disclose, inter alia, certain techniques whereby individualplayback devices configured to operate in a groupwise manner decide which playback device should function as a group coordinator for the group of playback devices.

[0149] U.S. App. 14 / 988,524, titled, “Multiple-Device Setup,” referred to as Docket No. 15-1103, filed Jan. 5, 2016, and issued May 28, 2019, as U.S. Pat. 10,303,422, discloses, inter alia, techniques for adding several new playback devices to a playback system at the same time, including, in scenarios when two or more of the same type of playback devices are detected during setup, causing one of the two or more playback devices to emit a sound that enables a user to identify the one playback device in connection with playback system configuration and setup.

[0150] U.S. App. 16 / 119,516, titled, “Media Playback System with Virtual Line-In,” referred to as Docket No. 18-0406, filed on Aug. 31, 2018, and issued on Oct. 22, 2019, as U.S. Pat. 10,452,345, and U.S. App. 16 / 119,642, titled, “Interoperability Of Native Media Playback System With Virtual Line-In,” referred to as Docket No. 18-0503, filed on Aug. 31, 2018, and issued on May 12, 2020, as U.S. Pat. 10,649,718, (both of which claim priority to U.S. Prov. App. 62 / 672,020, titled “Media Playback System with Virtual Line-In,” referred to as Docket No. 18-0406p, filed on May 15, 2018, and now expired) describe, inter alia, scenarios where one playback device in one playback system coordinates aspects of playback and control of another playback device in a different playback system in connection with facilitating interoperability between to two different playback systems.

[0151] U.S. App. 16 / 415,783, titled, “Wireless Multi-Channel Headphone Systems and Methods,” referred to as Docket No. 19-0303, filed on May 17, 2019, and issued on Nov. 16, 2021, as U.S. Pat. 11,178,504, describes, inter alia, a surround sound controller and one or more wireless headphones that switch between operating in various modes that have different latency characteristics. For example, in a first mode, the surround sound controller uses a first Modulation and Coding Scheme (MCS) to transmit first surround sound audio information to a first pair of headphones, and in a second mode, the surround sound controller uses a second MCS to transmit (a) the first surround sound audio information to the first pair of headphones and (b) second surround sound audio information to a second pair of headphones.

[0152] However, U.S. Apps. 13 / 489,674; 14 / 042,001; 14 / 041,989; 14 / 988,524; 16 / 119,516; 16 / 119,642; 62 / 672,020; and 16 / 415,783 do not describe the Area Zone configurations, Multi-Player Playback Devices, playback entity implementations, and / or thecalibration procedures described herein. The entire contents of U.S. Apps. 13 / 489,674; 14 / 042,001; 14 / 041,989; 14 / 988,524; 16 / 119,516; 16 / 119,642; 62 / 672,020; and 16 / 415,783 are incorporated herein by reference.

[0153] Further, some advancements over time have improved the sound quality of audio content played by playback devices via methods of calibrating playback settings (e.g., audio playback settings) for playback devices based on acoustic characteristics of the listening environment in which the playback devices are situated. At a high level, calibrating playback settings includes, inter alia, determining one or more acoustic characteristics of the listening environment, and adjusting one or more audio playback settings (e.g., equalization settings, relative loudness, playing timing delays, and / or perhaps other settings) based on the acoustic characteristics.

[0154] For example, U.S. App. 15 / 211,822, titled, “Spatial Audio Correction,” referred to as Docket No. 16-0402, filed on Jul. 15, 2016, and issued on Oct. 17, 2017, as U.S. Pat. 9,794,710 describes, inter alia, determining a spatial and / or spectral calibration for one or more playback devices within a listening area. Similarly, U.S. App. 16 / 115,524, titled “Playback Device Calibration,” referred to as Docket No. 18-0401, filed on Aug. 28, 2018, and issued on May 21, 2019, as U.S. Pat. 10,299,061 describes, inter alia, calibrating a playback device within a room so that the audio output by the playback device accounts for (e.g., offsets) acoustic characteristics of that room, thereby improving sound of the audio playback experienced by a listener within the room.

[0155] Additionally, U.S. App. 15 / 630,214, titled, “Immersive Audio in a Media Playback System,” referred to as Docket No. 16-0504, filed on Jun. 22, 2017, and issued on Jul. 17, 2018, as U.S. Pat. 10,028,069 describes, inter alia, processes that include obtaining audio responses from several different playback devices in a media playback system. For example, a first playback device at a first time plays back calibration audio while a microphone device records the calibration audio being played back. A second playback device at a second time plays back the calibration audio while the microphone device records the calibration audio being played. The process is repeated until every playback device in the media playback system has played the calibration audio and had its response recorded.

[0156] Further, IntT App. PCT / US22 / 77233, titled “Audio Parameter Adjustment Based on Playback Device Separation Distance,” referred to as Docket No. 21-0605-PCT, andpublished as WO 2023 / 056336 on Apr. 6, 2023, disclosed, inter alia, applying a low frequency filter to two or more devices in a zone based on a distance between the devices.

[0157] However, Apps. 15 / 211,822; 16 / 115,524; 15 / 630,214; and PCT7US22 / 77233 do not describe the Area Zone configurations, Multi-Player Playback Devices, playback entity implementations, and / or the calibration procedures described herein. The entire contents of Apps. 15 / 211,822; 16 / 115,524; 15 / 630,214; and PCT / US22 / 77233 are incorporated herein by reference.

[0158] While some examples described herein may refer to functions performed by given actors such as “users,” “listeners,” and / or other entities, it should be understood that this is for purposes of explanation only. The claims should not be interpreted to require action by any such example actor unless explicitly required by the language of the claims themselves.

[0159] In the Figures, identical reference numbers identify generally similar, and / or identical, elements. To facilitate the discussion of any particular element, the most significant digit or digits of a reference number refers to the Figure in which that element is first introduced. For example, element 110a is first introduced and discussed with reference to Figure 1 A. Many of the details, dimensions, angles and other features shown in the Figures are merely illustrative of particular example configurations of the disclosed technology. Accordingly, other example configurations can have other details, dimensions, angles and features without departing from the spirit or scope of the disclosure. In addition, those of ordinary skill in the art will appreciate that further example configurations of the various disclosed technologies can be practiced without several of the details described below.II. Suitable Operating Environment

[0160] Figure 1A is a partial cutaway view of a media playback system 100 distributed in an environment 101 (e.g., a house). The media playback system 100 comprises one or more playback devices 110 (identified individually as playback devices 1 lOa-n), one or more network microphone devices (“NMDs”), 120 (identified individually as NMDs 120a-c), and one or more control devices 130 (identified individually as control devices 130a and 130b).

[0161] As used herein the term “playback device” can generally refer to a network device configured to receive, process, and output data of a media playback system. For example, a playback device can be a network device that receives and processes audio content. In someexample configurations, a playback device includes one or more transducers or speakers powered by one or more amplifiers. In other example configurations, however, a playback device includes one of (or neither of) the speaker and the amplifier. For instance, a playback device can comprise one or more amplifiers configured to drive one or more speakers external to the playback device via a corresponding wire or cable.

[0162] Moreover, as used herein the term NMD (i.e., a “network microphone device”) can generally refer to a network device that is configured for audio detection. In some example configurations, an NMD is a stand-alone device configured primarily for audio detection. In other example configurations, an NMD is incorporated into a playback device (or vice versa).

[0163] The term “control device” can generally refer to a network device configured to perform functions relevant to facilitating user access, control, and / or configuration of the media playback system 100.

[0164] Each of the playback devices 110 is configured to receive audio signals or data from one or more media sources (e.g., one or more remote servers, one or more local devices) and play back the received audio signals or data as sound. The one or more NMDs 120 are configured to receive spoken word commands, and the one or more control devices 130 are configured to receive user input. In response to the received spoken word commands and / or user input, the media playback system 100 can play back audio via one or more of the playback devices 110. In certain example configurations, the playback devices 110 are configured to commence playback of media content in response to a trigger. For instance, one or more of the playback devices 110 can be configured to play back a morning playlist upon detection of an associated trigger condition (e.g., presence of a user in a kitchen, detection of a coffee machine operation). In some example configurations, for example, the media playback system 100 is configured to play back audio from a first playback device (e.g., the playback device 100a) in synchrony with a second playback device (e.g., the playback device 100b). Interactions between the playback devices 110, NMDs 120, and / or control devices 130 of the media playback system 100 configured in accordance with the various example configurations of the disclosure are described in greater detail below with respect to Figures 1B-1L.

[0165] In the illustrated embodiment of Figure 1A, the environment 101 comprises a household having several rooms, spaces, and / or playback zones, including (clockwise from upper left) a master bathroom 101a, a master bedroom 101b, a second bedroom 101c, a familyroom or den 101 d, an office lOle, a living room l Olf, a dining room 101g, a kitchen lOlh, and an outdoor patio lOli. While certain example configurations and examples are described below in the context of a home environment, the technologies described herein may be implemented in other types of environments. In some example configurations, for example, the media playback system 100 can be implemented in one or more commercial settings (e.g., a restaurant, mall, airport, hotel, a retail or other store), one or more vehicles (e.g., a sports utility vehicle, bus, car, a ship, a boat, an airplane), multiple environments (e.g., a combination of home and vehicle environments), and / or another suitable environment where multi-zone audio may be desirable.

[0166] The media playback system 100 can comprise one or more playback zones, some of which may correspond to the rooms in the environment 101. The media playback system 100 can be established with one or more playback zones, after which additional zones may be added, or removed to form, for example, the configuration shown in Figure 1 A. Each zone may be given a name according to a different room or space such as the office lOle, master bathroom 101a, master bedroom 101b, the second bedroom 101c, kitchen lOlh, dining room 101g, living room 10 If, and / or the outdoor patio lOli. In some aspects, a single playback zone may include multiple rooms or spaces. In certain aspects, a single room or space may include multiple playback zones.

[0167] In the illustrated embodiment of Figure 1A, the master bathroom 101a, the second bedroom 101c, the office lOle, the living room 10 If, the dining room 101g, the kitchen lOlh, and the outdoor patio lOli each include one playback device 110, and the master bedroom 101b and the den 101 d include a plurality of playback devices 110. In the master bedroom 101b, the playback devices 1101 and 110m may be configured, for example, to play back audio content in synchrony as individual ones of playback devices 110, as a bonded playback zone, as a consolidated playback device, and / or any combination thereof. Similarly, in the den 10 Id, the playback devices 1 lOh-j can be configured, for instance, to play back audio content in synchrony as individual ones of playback devices 110, as one or more bonded playback devices, and / or as one or more consolidated playback devices. Additional details regarding bonded and consolidated playback devices are described below with respect to, for example, Figures IB and IE and 1I-1M.

[0168] In some aspects, one or more of the playback zones in the environment 101 may each be playing different audio content. For instance, a user may be grilling on the patio lOli andlistening to hip hop music being played by the playback device 110c while another user is preparing food in the kitchen lOlh and listening to classical music played by the playback device 110b. In another example, a playback zone may play the same audio content in synchrony with another playback zone. For instance, the user may be in the office lOle listening to the playback device 1 lOf playing back the same hip hop music being played back by playback device 110c on the patio lOli. In some aspects, the playback devices 110c and 1 lOf play back the hip hop music in synchrony such that the user perceives that the audio content is being played seamlessly (or at least substantially seamlessly) while moving between different playback zones. Additional details regarding audio playback synchronization among playback devices and / or zones can be found, for example, in U.S. Patent No. 8,234,395 entitled, “System and method for synchronizing operations among a plurality of independently clocked digital data processing devices,” which is incorporated herein by reference in its entirety. a. Suitable Media Playback System

[0169] Figure IB is a schematic diagram of the media playback system 100 and a cloud network 102. For ease of illustration, certain devices of the media playback system 100 and the cloud network 102 are omitted from Figure IB. One or more communications links 103 (referred to hereinafter as “the links 103”) communicatively couple the media playback system 100 and the cloud network 102.

[0170] The links 103 can comprise, for example, one or more wired networks, one or more wireless networks, one or more wide area networks (WAN), one or more local area networks (LAN), one or more personal area networks (PAN), one or more telecommunication networks (e.g., one or more Global System for Mobiles (GSM) networks, Code Division Multiple Access (CDMA) networks, Long-Term Evolution (LTE) networks, 5G communication network networks, and / or other suitable data transmission protocol networks), etc. The cloud network 102 is configured to deliver media content (e.g., audio content, video content, photographs, social media content) to the media playback system 100 in response to a request transmitted from the media playback system 100 via the links 103. In some example configurations, the cloud network 102 is further configured to receive data (e.g. voice input data) from the media playback system 100 and correspondingly transmit commands and / or media content to the media playback system 100.

[0171] The cloud network 102 comprises computing devices 106 (identified separately as a first computing device 106a, a second computing device 106b, and a third computing device 106c). The computing devices 106 can comprise individual computers or servers, such as, for example, a media streaming service server storing audio and / or other media content, a voice service server, a social media server, a media playback system control server, etc. In some example configurations, one or more of the computing devices 106 comprise modules of a single computer or server. In certain example configurations, one or more of the computing devices 106 comprise one or more modules, computers, and / or servers. Moreover, while the cloud network 102 is described above in the context of a single cloud network, in some example configurations the cloud network 102 comprises a plurality of cloud networks comprising communicatively coupled computing devices. Furthermore, while the cloud network 102 is shown in Figure IB as having three of the computing devices 106, in some example configurations, the cloud network 102 comprises fewer (or more than) three computing devices 106.

[0172] The media playback system 100 is configured to receive media content from the networks 102 via the links 103. The received media content can comprise, for example, a Uniform Resource Identifier (URI) and / or a Uniform Resource Locator (URL). For instance, in some examples, the media playback system 100 can stream, download, or otherwise obtain data from a URI or a URL corresponding to the received media content. A network 104 communicatively couples the links 103 and at least a portion of the devices (e.g., one or more of the playback devices 110, NMDs 120, and / or control devices 130) of the media playback system 100. The network 104 can include, for example, a wireless network (e.g., a WiFi network, a Bluetooth, a Z-Wave network, a ZigBee, and / or other suitable wireless communication protocol network) and / or a wired network (e.g., a network comprising Ethernet, Universal Serial Bus (USB), and / or another suitable wired communication). As those of ordinary skill in the art will appreciate, as used herein, “WiFi” can refer to several different communication protocols including, for example, Institute of Electrical and Electronics Engineers (IEEE) 802.11a, 802.11b, 802.11g, 802.1 In, 802.1 lac, 802.1 lac, 802. Had, 802.11af, 802. Hah, 802.1 lai,802.1 laj, 802. Haq, 802.11ax, 802. Hay, 802.15, etc. transmitted at 2.4 Gigahertz (GHz), 5 GHz, and / or another suitable frequency.

[0173] In some example configurations, the network 104 comprises a dedicated communication network that the media playback system 100 uses to transmit messages betweenindividual devices and / or to transmit media content to and from media content sources (e.g., one or more of the computing devices 106). In certain example configurations, the network 104 is configured to be accessible only to devices in the media playback system 100, thereby reducing interference and competition with other household devices. In other example configurations, however, the network 104 comprises an existing household communication network (e.g., a household WiFi network). In some example configurations, the links 103 and the network 104 comprise one or more of the same networks. In some aspects, for example, the links 103 and the network 104 comprise a telecommunication network (e.g., an LTE network, a 5G network). Moreover, in some example configurations, the media playback system 100 is implemented without the network 104, and devices comprising the media playback system 100 can communicate with each other, for example, via one or more direct connections, PANs, telecommunication networks, and / or other suitable communications links.

[0174] In some example configurations, audio content sources may be regularly added or removed from the media playback system 100. In some example configurations, for example, the media playback system 100 performs an indexing of media items when one or more media content sources are updated, added to, and / or removed from the media playback system 100. The media playback system 100 can scan identifiable media items in some or all folders and / or directories accessible to the playback devices 110, and generate or update a media content database comprising metadata (e.g., title, artist, album, track length) and other associated information (e.g., URIs, URLs) for each identifiable media item found. In some example configurations, for example, the media content database is stored on one or more of the playback devices 110, network microphone devices 120, and / or control devices 130.

[0175] In the illustrated embodiment of Figure IB, the playback devices 1101 and 110m comprise a group 107a. The playback devices 1101 and 110m can be positioned in different rooms in a household and be grouped together in the group 107a on a temporary or permanent basis based on user input received at the control device 130a and / or another control device 130 in the media playback system 100. When arranged in the group 107a, the playback devices 1101 and 110m can be configured to play back the same or similar audio content in synchrony from one or more audio content sources. In certain example configurations, for example, the group 107a comprises a bonded zone in which the playback devices 1101 and 110m comprise left audio and right audio channels, respectively, of multi-channel audio content, thereby producing orenhancing a stereo effect of the audio content. In some example configurations, the group 107a includes additional playback devices 110. In other example configurations, however, the media playback system 100 omits the group 107a and / or other grouped arrangements of the playback devices 110. Additional details regarding groups and other arrangements of playback devices are described in further detail below with respect to Figures 1-1 through IM.

[0176] The media playback system 100 includes the NMDs 120a and 120d, each comprising one or more microphones configured to receive voice utterances from a user. In the illustrated embodiment of Figure IB, the NMD 120a is a standalone device and the NMD 120d is integrated into the playback device 1 lOn. The NMD 120a, for example, is configured to receive voice input 121 from a user 123. In some example configurations, the NMD 120a transmits data associated with the received voice input 121 to a voice assistant service (VAS) configured to (i) process the received voice input data and (ii) transmit a corresponding command to the media playback system 100. In some aspects, for example, the computing device 106c comprises one or more modules and / or servers of a VAS (e.g., a VAS operated by one or more of SONOS®, AMAZON®, GOOGLE® APPLE®, MICROSOFT®). The computing device 106c can receive the voice input data from the NMD 120a via the network 104 and the links 103. In response to receiving the voice input data, the computing device 106c processes the voice input data (i.e., “Play Hey Jude by The Beatles”), and determines that the processed voice input includes a command to play a song (e.g., “Hey Jude”). The computing device 106c accordingly transmits commands to the media playback system 100 to play back “Hey Jude” by the Beatles from a suitable media service (e.g., via one or more of the computing devices 106) on one or more of the playback devices 110. b. Suitable Playback Devices

[0177] Figure 1C is a block diagram of the playback device 110a comprising an input / output 111. The input / output 111 can include an analog EO I l la (e.g., one or more wires, cables, and / or other suitable communications links configured to carry analog signals) and / or a digital EO 11 lb (e.g., one or more wires, cables, or other suitable communications links configured to carry digital signals). In some example configurations, the analog EO 11 la is an audio line-in input connection comprising, for example, an auto-detecting 3.5mm audio line-in connection. In some example configurations, the digital EO 111b comprises a Sony / PhilipsDigital Interface Format (S / PDIF) communication interface and / or cable and / or a Toshiba Link (TOSLINK) cable. In some example configurations, the digital VO 111b comprises an High- Definition Multimedia Interface (HDMI) interface and / or cable. In some example configurations, the digital VO 111b includes one or more wireless communications links comprising, for example, a radio frequency (RF), infrared, WiFi, Bluetooth, or another suitable communication protocol. In certain example configurations, the analog VO 1 I la and the digital VO 11 lb comprise interfaces (e.g., ports, plugs, jacks) configured to receive connectors of cables transmitting analog and digital signals, respectively, without necessarily including cables.

[0178] The playback device 110a, for example, can receive media content (e.g., audio content comprising music and / or other sounds) from a local audio source 105 via the input / output 111 (e.g., a cable, a wire, a PAN, a Bluetooth connection, an ad hoc wired or wireless communication network, and / or another suitable communications link). The local audio source 105 can comprise, for example, a mobile device (e.g., a smartphone, a tablet, a laptop computer) or another suitable audio component (e.g., a television, a desktop computer, an amplifier, a phonograph, a Blu-ray player, a memory storing digital media files). In some aspects, the local audio source 105 includes local music libraries on a smartphone, a computer, a networked-attached storage (NAS), and / or another suitable device configured to store media files. In certain example configurations, one or more of the playback devices 110, NMDs 120, and / or control devices 130 comprise the local audio source 105. In other example configurations, however, the media playback system omits the local audio source 105 altogether. In some example configurations, the playback device 110a does not include an input / output 111 and receives all audio content via the network 104.

[0179] The playback device 110a further comprises electronics 112, a user interface 113 (e.g., one or more buttons, knobs, dials, touch- sensitive surfaces, displays, touchscreens), and one or more transducers 114 (referred to hereinafter as “the transducers 114”). The electronics 112 is configured to receive audio from an audio source (e.g., the local audio source 105) via the input / output 111, one or more of the computing devices 106a-c via the network 104 (Figure IB)), amplify the received audio, and output the amplified audio for playback via one or more of the transducers 114. In some example configurations, the playback device 110a optionally includes one or more microphones 115 (e.g., a single microphone, a plurality of microphones, a microphone array) (hereinafter referred to as “the microphones 115”). In certain exampleconfigurations, for example, the playback device 110a having one or more of the optional microphones 115 can operate as an NMD configured to receive voice input from a user and correspondingly perform one or more operations based on the received voice input.

[0180] In the illustrated embodiment of Figure 1C, the electronics 112 comprise one or more processors 112a (referred to hereinafter as “the processors 112a”), memory 112b, software components 112c, a network interface 112d, one or more audio processing components 112g (referred to hereinafter as “the audio components 112g”), one or more audio amplifiers 112h (referred to hereinafter as “the amplifiers 112h”), and power 112i (e.g., one or more power supplies, power cables, power receptacles, batteries, induction coils, Power-over Ethernet (POE) interfaces, and / or other suitable sources of electric power). In some example configurations, the electronics 112 optionally include one or more other components 112j (e.g., one or more sensors, video displays, touchscreens, battery charging bases).

[0181] The processors 112a can comprise clock-driven computing component(s) configured to process data, and the memory 112b can comprise a computer-readable medium (e.g., a tangible, non-transitory computer-readable medium, data storage loaded with one or more of the software components 112c) configured to store instructions for performing various operations and / or functions. The processors 112a are configured to execute the instructions stored on the memory 112b to perform one or more of the operations. The operations can include, for example, causing the playback device 110a to retrieve audio information from an audio source (e.g., one or more of the computing devices 106a-c (Figure IB)), and / or another one of the playback devices 110. In some example configurations, the operations further include causing the playback device 110a to send audio information to another one of the playback devices 110a and / or another device (e.g., one of the NMDs 120). Certain example configurations include operations causing the playback device 110a to pair with another of the one or more playback devices 110 to enable a multi-channel audio environment (e.g., a stereo pair, a bonded zone).

[0182] The processors 112a can be further configured to perform operations causing the playback device 110a to synchronize playback of audio content with another of the one or more playback devices 110. As those of ordinary skill in the art will appreciate, during synchronous playback of audio content on a plurality of playback devices, a listener will preferably be unable to perceive time-delay differences between playback of the audio content by the playback device110a and the other one or more other playback devices 1 10. Additional details regarding audio playback synchronization among playback devices can be found, for example, in U.S. Patent No. 8,234,395, which was incorporated by reference above.

[0183] In some example configurations, the memory 112b is further configured to store data associated with the playback device 110a, such as one or more zones and / or zone groups of which the playback device 110a is a member, audio sources accessible to the playback device 110a, and / or a playback queue that the playback device 110a (and / or another of the one or more playback devices) can be associated with. The stored data can comprise one or more state variables that are periodically updated and used to describe a state of the playback device 110a. The memory 112b can also include data associated with a state of one or more of the other devices (e.g., the playback devices 110, NMDs 120, control devices 130) of the media playback system 100. In some aspects, for example, the state data is shared during predetermined intervals of time (e.g., every 5 seconds, every 10 seconds, every 60 seconds) among at least a portion of the devices of the media playback system 100, so that one or more of the devices have the most recent data associated with the media playback system 100.

[0184] The network interface 112d is configured to facilitate a transmission of data between the playback device 110a and one or more other devices on a data network such as, for example, the links 103 and / or the network 104 (Figure IB). The network interface 112d is configured to transmit and receive data corresponding to media content (e.g., audio content, video content, text, photographs) and other signals (e.g., non-transitory signals) comprising digital packet data including an Internet Protocol (IP)-based source address and / or an IP-based destination address. The network interface 112d can parse the digital packet data such that the electronics 112 properly receives and processes the data destined for the playback device 110a.

[0185] In the illustrated embodiment of Figure 1C, the network interface 112d comprises one or more wireless interfaces 112e (referred to hereinafter as “the wireless interface 112e”). The wireless interface 112e (e.g., a suitable interface comprising one or more antennae) can be configured to wirelessly communicate with one or more other devices (e.g., one or more of the other playback devices 110, NMDs 120, and / or control devices 130) that are communicatively coupled to the network 104 (Figure IB) in accordance with a suitable wireless communication protocol (e.g., WiFi, Bluetooth, LTE). In some example configurations, the network interface 112d optionally includes a wired interface 112f (e.g., an interface or receptacle configured toreceive a network cable such as an Ethernet, a USB-A, USB-C, and / or Thunderbolt cable) configured to communicate over a wired connection with other devices in accordance with a suitable wired communication protocol. In certain example configurations, the network interface 112d includes the wired interface 112f and excludes the wireless interface 112e. In some example configurations, the electronics 112 excludes the network interface 112d altogether and transmits and receives media content and / or other data via another communication path (e.g., the input / output 111).

[0186] The audio processing components 112g are configured to process and / or filter data comprising media content received by the electronics 112 (e.g., via the input / output 111 and / or the network interface 112d) to produce output audio signals. In some example configurations, the audio processing components 112g comprise, for example, one or more digital-to-analog converters (DAC), audio preprocessing components, audio enhancement components, a digital signal processors (DSPs), and / or other suitable audio processing components, modules, circuits, etc. In certain example configurations, one or more of the audio processing components 112g can comprise one or more subcomponents of the processors 112a. In some example configurations, the electronics 112 omits the audio processing components 112g. In some aspects, for example, the processors 112a execute instructions stored on the memory 112b to perform audio processing operations to produce the output audio signals.

[0187] The amplifiers 112h are configured to receive and amplify the audio output signals produced by the audio processing components 112g and / or the processors 112a. The amplifiers 112h can comprise electronic devices and / or components configured to amplify audio signals to levels sufficient for driving one or more of the transducers 114. In some example configurations, for example, the amplifiers 112h include one or more switching or class-D power amplifiers. In other example configurations, however, the amplifiers include one or more other types of power amplifiers (e.g., linear gain power amplifiers, class-A amplifiers, class-B amplifiers, class-AB amplifiers, class-C amplifiers, class-D amplifiers, class-E amplifiers, class- F amplifiers, class-G and / or class H amplifiers, and / or another suitable type of power amplifier). In certain example configurations, the amplifiers 112h comprise a suitable combination of two or more of the foregoing types of power amplifiers. Moreover, in some example configurations, individual ones of the amplifiers 112h correspond to individual ones of the transducers 114. In other example configurations, however, the electronics 112 includes a single one of theamplifiers 112h configured to output amplified audio signals to a plurality of the transducers 114. In some other example configurations, the electronics 112 omits the amplifiers 112h.

[0188] The transducers 114 (e.g., one or more speakers and / or speaker drivers) receive the amplified audio signals from the amplifier 112h and render or output the amplified audio signals as sound (e.g., audible sound waves having a frequency between about 20 Hertz (Hz) and 20 kilohertz (kHz)). In some example configurations, the transducers 114 can comprise a single transducer. In other example configurations, however, the transducers 114 comprise a plurality of audio transducers. In some example configurations, the transducers 114 comprise more than one type of transducer. For example, the transducers 114 can include one or more low frequency transducers (e.g., subwoofers, woofers), mid-range frequency transducers (e.g., mid-range transducers, mid-woofers), and one or more high frequency transducers (e.g., one or more tweeters). As used herein, “low frequency” can generally refer to audible frequencies below about 500 Hz, “mid-range frequency” can generally refer to audible frequencies between about 500 Hz and about 2 kHz, and “high frequency” can generally refer to audible frequencies above 2 kHz. In certain example configurations, however, one or more of the transducers 114 comprise transducers that do not adhere to the foregoing frequency ranges. For example, one of the transducers 114 may comprise a mid-woofer transducer configured to output sound at frequencies between about 200 Hz and about 5 kHz.

[0189] By way of illustration, SONOS, Inc. presently offers (or has offered) for sale certain playback devices including, for example, a “SONOS ONE,” “PLAY:1,” “PLAY:3,” “PLAYA,” “PLAYBAR,” “PLAYBASE,” “CONNECT: AMP,” “CONNECT,” and “SUB.” Other suitable playback devices may additionally or alternatively be used to implement the playback devices of example configurations disclosed herein. Additionally, one of ordinary skilled in the art will appreciate that a playback device is not limited to the examples described herein or to SONOS product offerings. In some example configurations, for example, one or more playback devices 110 comprises wired or wireless headphones (e.g., over-the-ear headphones, on-ear headphones, in-ear earphones). In other example configurations, one or more of the playback devices 110 comprise a docking station and / or an interface configured to interact with a docking station for personal mobile media playback devices. In certain example configurations, a playback device may be integral to another device or component such as a television, a lighting fixture, or some other device for indoor or outdoor use. In some exampleconfigurations, a playback device omits a user interface and / or one or more transducers. For example, FIG. ID is a block diagram of a playback device 1 lOp comprising the input / output 111 and electronics 112 without the user interface 113 or transducers 114.

[0190] Figure IE is a block diagram of a bonded playback device 1 lOq comprising the playback device 110a (Figure 1C) sonically bonded with the playback device 1 lOi (e.g., a subwoofer) (Figure 1A). In the illustrated embodiment, the playback devices 110a and 1 lOi are separate ones of the playback devices 110 housed in separate enclosures. In some example configurations, however, the bonded playback device HOq comprises a single enclosure housing both the playback devices 110a and 1 lOi. The bonded playback device 1 lOq can be configured to process and reproduce sound differently than an unbonded playback device (e.g., the playback device 110a of Figure 1C) and / or paired or bonded playback devices (e.g., the playback devices 1101 and 110m of Figure IB). In some example configurations, for example, the playback device 110a is full-range playback device configured to render low frequency, mid-range frequency, and high frequency audio content, and the playback device 1 lOi is a subwoofer configured to render low frequency audio content. In some aspects, the playback device 110a, when bonded with the first playback device, is configured to render only the mid-range and high frequency components of a particular audio content, while the playback device 1 lOi renders the low frequency component of the particular audio content. In some example configurations, the bonded playback device 1 lOq includes additional playback devices and / or another bonded playback device.Additional playback device example configurations are described in further detail below with respect to Figures 2A-3D. c. Suitable Network Microphone Devices (NMDs)

[0191] Figure IF is a block diagram of the NMD 120a (Figures 1 A and IB). The NMD 120a includes one or more voice processing components 124 (hereinafter “the voice components 124”) and several components described with respect to the playback device 110a (Figure 1C) including the processors 112a, the memory 112b, and the microphones 115. The NMD 120a optionally comprises other components also included in the playback device 110a (Figure 1C), such as the user interface 113 and / or the transducers 114. In some example configurations, the NMD 120a is configured as a media playback device (e.g., one or more of the playback devices 110), and further includes, for example, one or more of the audio processing components 112g(Figure 1 C), the transducers 114, and / or other playback device components. In certain example configurations, the NMD 120a comprises an Internet of Things (loT) device such as, for example, a thermostat, alarm panel, fire and / or smoke detector, etc. In some example configurations, the NMD 120a comprises the microphones 115, the voice processing 124, and only a portion of the components of the electronics 112 described above with respect to Figure IB. In some aspects, for example, the NMD 120a includes the processor 112a and the memory 112b (Figure IB), while omitting one or more other components of the electronics 112. In some example configurations, the NMD 120a includes additional components (e.g., one or more sensors, cameras, thermometers, barometers, hygrometers).

[0192] In some example configurations, an NMD can be integrated into a playback device. Figure 1G is a block diagram of a playback device 1 lOr comprising an NMD 120d. The playback device 1 lOr can comprise many or all of the components of the playback device 110a and further include the microphones 115 and voice processing 124 (Figure IF). The playback device 1 lOr optionally includes an integrated control device 130c. The control device 130c can comprise, for example, a user interface (e.g., the user interface 113 of Figure IB) configured to receive user input (e.g., touch input, voice input) without a separate control device. In other example configurations, however, the playback device 1 lOr receives commands from another control device (e.g., the control device 130a of Figure IB). Additional NMD example configurations are described in further detail below with respect to Figures 3A-3F.

[0193] Referring again to Figure IF, the microphones 115 are configured to acquire, capture, and / or receive sound from an environment (e.g., the environment 101 of Figure 1A) and / or a room in which the NMD 120a is positioned. The received sound can include, for example, vocal utterances, audio played back by the NMD 120a and / or another playback device, background voices, ambient sounds, etc. The microphones 115 convert the received sound into electrical signals to produce microphone data. The voice processing 124 receives and analyzes the microphone data to determine whether a voice input is present in the microphone data. The voice input can comprise, for example, an activation word followed by an utterance including a user request. As those of ordinary skill in the art will appreciate, an activation word is a word or other audio cue that signifying a user voice input. For instance, in querying the AMAZON® VAS, a user might speak the activation word "Alexa." Other examples include "Ok, Google" for invoking the GOOGLE® VAS and "Hey, Siri" for invoking the APPLE® VAS.

[0194] After detecting the activation word, voice processing 124 monitors the microphone data for an accompanying user request in the voice input. The user request may include, for example, a command to control a third-party device, such as a thermostat (e.g., NEST® thermostat), an illumination device (e.g., a PHILIPS HUE ® lighting device), or a media playback device (e.g., a Sonos® playback device). For example, a user might speak the activation word “Alexa” followed by the utterance “set the thermostat to 68 degrees” to set a temperature in a home (e.g., the environment 101 of Figure 1A). The user might speak the same activation word followed by the utterance “turn on the living room” to turn on illumination devices in a living room area of the home. The user may similarly speak an activation word followed by a request to play a particular song, an album, or a playlist of music on a playback device in the home. Additional description regarding receiving and processing voice input data can be found in further detail below with respect to Figures 3 A-3F. d. Suitable Control Devices

[0195] Figure 1H is a partially schematic diagram of the control device 130a (Figures 1 A and IB). As used herein, the term “control device” can be used interchangeably with “controller” or “control system.” Among other features, the control device 130a is configured to receive user input related to the media playback system 100 and, in response, cause one or more devices in the media playback system 100 to perform an action(s) or operation(s) corresponding to the user input. In the illustrated embodiment, the control device 130a comprises a smartphone (e.g., an iPhone™ an Android phone) on which media playback system controller application software is installed. In some example configurations, the control device 130a comprises, for example, a tablet (e.g., an iPad™), a computer (e.g., a laptop computer, a desktop computer), and / or another suitable device (e.g., a television, an automobile audio head unit, an loT device). In certain example configurations, the control device 130a comprises a dedicated controller for the media playback system 100. In other example configurations, as described above with respect to Figure 1G, the control device 130a is integrated into another device in the media playback system 100 (e.g., one more of the playback devices 110, NMDs 120, and / or other suitable devices configured to communicate over a network).

[0196] The control device 130a includes electronics 132, a user interface 133, one or more speakers 134, and one or more microphones 135. The electronics 132 comprise one ormore processors 132a (referred to hereinafter as “the processors 132a”), a memory 132b, software components 132c, and a network interface 132d. The processor 132a can be configured to perform functions relevant to facilitating user access, control, and configuration of the media playback system 100. The memory 132b can comprise data storage that can be loaded with one or more of the software components executable by the processor 302 to perform those functions. The software components 132c can comprise applications and / or other executable software configured to facilitate control of the media playback system 100. The memory 112b can be configured to store, for example, the software components 132c, media playback system controller application software, and / or other data associated with the media playback system 100 and the user.

[0197] The network interface 132d is configured to facilitate network communications between the control device 130a and one or more other devices in the media playback system 100, and / or one or more remote devices. In some example configurations, the network interface 132d is configured to operate according to one or more suitable communication industry standards (e.g., infrared, radio, wired standards including IEEE 802.3, wireless standards including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G, LTE). The network interface 132d can be configured, for example, to transmit data to and / or receive data from the playback devices 110, the NMDs 120, other ones of the control devices 130, one of the computing devices 106 of Figure IB, devices comprising one or more other media playback systems, etc. The transmitted and / or received data can include, for example, playback device control commands, state variables, playback zone and / or zone group configurations. For instance, based on user input received at the user interface 133, the network interface 132d can transmit a playback device control command (e.g., volume control, audio playback control, audio content selection) from the control device 304 to one or more of playback devices. The network interface 132d can also transmit and / or receive configuration changes such as, for example, adding / removing one or more playback devices to / from a zone, adding / removing one or more zones to / from a zone group, forming a bonded or consolidated player, separating one or more playback devices from a bonded or consolidated player, among others. Additional description of zones and groups can be found below with respect to Figures 1-1 through IM.

[0198] The user interface 133 is configured to receive user input and can facilitate 'control of the media playback system 100. The user interface 133 includes media content art133a (e g., album art, lyrics, videos), a playback status indicator 133b (e g., an elapsed and / or remaining time indicator), media content information region 133c, a playback control region 133d, and a zone indicator 133e. The media content information region 133c can include a display of relevant information (e.g., title, artist, album, genre, release year) about media content currently playing and / or media content in a queue or playlist. The playback control region 133d can include selectable (e.g., via touch input and / or via a cursor or another suitable selector) icons to cause one or more playback devices in a selected playback zone or zone group to perform playback actions such as, for example, play or pause, fast forward, rewind, skip to next, skip to previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit cross fade mode, etc. The playback control region 133d may also include selectable icons to modify equalization settings, playback volume, and / or other suitable playback actions. In the illustrated embodiment, the user interface 133 comprises a display presented on a touch screen interface of a smartphone (e.g., an iPhone™ an Android phone). In some example configurations, however, user interfaces of varying formats, styles, and interactive sequences may alternatively be implemented on one or more network devices to provide comparable control access to a media playback system.

[0199] The one or more speakers 134 (e.g., one or more transducers) can be configured to output sound to the user of the control device 130a. In some example configurations, the one or more speakers comprise individual transducers configured to correspondingly output low frequencies, mid-range frequencies, and / or high frequencies. In some aspects, for example, the control device 130a is configured as a playback device (e.g., one of the playback devices 110). Similarly, in some example configurations the control device 130a is configured as an NMD (e.g., one of the NMDs 120), receiving voice commands and other sounds via the one or more microphones 135.

[0200] The one or more microphones 135 can comprise, for example, one or more condenser microphones, electret condenser microphones, dynamic microphones, and / or other suitable types of microphones or transducers. In some example configurations, two or more of the microphones 135 are arranged to capture location information of an audio source (e.g., voice, audible sound) and / or configured to facilitate filtering of background noise. Moreover, in certain example configurations, the control device 130a is configured to operate as playback device and an NMD. In other example configurations, however, the control device 130a omits the one or more speakers 134 and / or the one or more microphones 135. For instance, the control device130a may comprise a device (e.g., a thermostat, an loT device, a network device) comprising a portion of the electronics 132 and the user interface 133 (e.g., a touch screen) without any speakers or microphones. Additional control device example configurations are described in further detail below with respect to Figures 4A-4D and 5. e. Suitable Playback Device Configurations

[0201] Figures 1-1 through IM show example configurations of playback devices in zones and zone groups. Referring first to Figure IM, in one example, a single playback device may belong to a zone. For example, the playback device 110g in the second bedroom 101c (FIG. 1A) may belong to Zone C. In some implementations described below, multiple playback devices may be “bonded” to form a “bonded pair” which together form a single zone. For example, the playback device 1101 (e.g., a left playback device) can be bonded to the playback device 1101 (e.g., a left playback device) to form Zone A. Bonded playback devices may have different playback responsibilities (e.g., channel responsibilities). In another implementation described below, multiple playback devices may be merged to form a single zone. For example, the playback device 1 lOh (e.g., a front playback device) may be merged with the playback device 1 lOi (e.g., a subwoofer), and the playback devices 1 lOj and 110k (e.g., left and right surround speakers, respectively) to form a single Zone D. In another example, the playback devices 110g and 1 lOh can be merged to form a merged group or a zone group 108b. The merged playback devices 110g and 1 lOh may not be specifically assigned different playback responsibilities. That is, the merged playback devices 1 lOh and 1 lOi may, aside from playing audio content in synchrony, each play audio content as they would if they were not merged.

[0202] Each zone in the media playback system 100 may be provided for control as a single user interface (UI) entity. For example, Zone A may be provided as a single entity named Master Bathroom. Zone B may be provided as a single entity named Master Bedroom. Zone C may be provided as a single entity named Second Bedroom.

[0203] Playback devices that are bonded may have different playback responsibilities, such as responsibilities for certain audio channels. For example, as shown in Figure 1-1, the playback devices 1101 and 110m may be bonded so as to produce or enhance a stereo effect of audio content. In this example, the playback device 1101 may be configured to play a left channel audio component, while the playback device 110k may be configured to play a right channelaudio component. In some implementations, such stereo bonding may be referred to as “pairing.”

[0204] Additionally, bonded playback devices may have additional and / or different respective speaker drivers. As shown in Figure 1J, the playback device 1 lOh named Front may be bonded with the playback device 1 lOi named SUB. The Front device 1 lOh can be configured to render a range of mid to high frequencies and the SUB device 1 lOi can be configured render low frequencies. When unbonded, however, the Front device 1 lOh can be configured render a full range of frequencies. As another example, Figure IK shows the Front and SUB devices I lOh and 1 lOi further bonded with Left and Right playback devices 1 lOj and 110k, respectively. In some implementations, the Right and Left devices 1 lOj and 102k can be configured to form surround or “satellite” channels of a home theater system. The bonded playback devices 1 lOh, 1 lOi, 1 lOj, and 110k may form a single Zone D (FIG. IM).

[0205] Playback devices that are merged may not have assigned playback responsibilities, and may each render the full range of audio content the respective playback device is capable of. Nevertheless, merged devices may be represented as a single UI entity (i.e., a zone, as discussed above). For instance, the playback devices 110a and 1 lOn the master bathroom have the single UI entity of Zone A. In one embodiment, the playback devices 110a and 1 lOn may each output the full range of audio content each respective playback devices 110a and 1 lOn are capable of, in synchrony.

[0206] In some example configurations, an NMD is bonded or merged with another device so as to form a zone. For example, the NMD 120b may be bonded with the playback device 1 lOe, which together form Zone F, named Living Room. In other example configurations, a stand-alone network microphone device may be in a zone by itself. In other example configurations, however, a stand-alone network microphone device may not be associated with a zone. Additional details regarding associating network microphone devices and playback devices as designated or default devices may be found, for example, in previously referenced U.S. Patent Application No. 15 / 438,749.

[0207] Zones of individual, bonded, and / or merged devices may be grouped to form a zone group. For example, referring to Figure IM, Zone A may be grouped with Zone B to form a zone group 108a that includes the two zones. Similarly, Zone G may be grouped with Zone H to form the zone group 108b. As another example, Zone A may be grouped with one or moreother Zones C-I. The Zones A-I may be grouped and ungrouped in numerous ways. For example, three, four, five, or more (e.g., all) of the Zones A-I may be grouped. When grouped, the zones of individual and / or bonded playback devices may play back audio in synchrony with one another, as described in previously referenced U.S. Patent No. 8,234,395. Playback devices may be dynamically grouped and ungrouped to form new or different groups that synchronously play back audio content.

[0208] In various implementations, the zones in an environment may be the default name of a zone within the group or a combination of the names of the zones within a zone group. For example, Zone Group 108b can have be assigned a name such as “Dining + Kitchen”, as shown in Figure IM. In some example configurations, a zone group may be given a unique name selected by a user.

[0209] Certain data may be stored in a memory of a playback device (e.g., the memory 112b of Figure 1C) as one or more state variables that are periodically updated and used to describe the state of a playback zone, the playback device(s), and / or a zone group associated therewith. The memory may also include the data associated with the state of the other devices of the media system, and shared from time to time among the devices so that one or more of the devices have the most recent data associated with the system.

[0210] In some example configurations, the memory may store instances of various variable types associated with the states. Variables instances may be stored with identifiers (e.g., tags) corresponding to type. For example, certain identifiers may be a first type “al” to identify playback device(s) of a zone, a second type “bl” to identify playback device(s) that may be bonded in the zone, and a third type “cl” to identify a zone group to which the zone may belong. As a related example, identifiers associated with the second bedroom 101c may indicate that the playback device is the only playback device of the Zone C and not in a zone group. Identifiers associated with the Den may indicate that the Den is not grouped with other zones but includes bonded playback devices 11 Oh- 110k. Identifiers associated with the Dining Room may indicate that the Dining Room is part of the Dining + Kitchen zone group 108b and that devices 110b and 1 lOd are grouped (FIG. IL). Identifiers associated with the Kitchen may indicate the same or similar information by virtue of the Kitchen being part of the Dining + Kitchen zone group 108b. Other example zone variables and identifiers are described below.

[0211] In yet another example, the media playback system 100 may variables or identifiers representing other associations of zones and zone groups, such as identifiers associated with Areas, as shown in Figure IM. An area may involve a cluster of zone groups and / or zones not within a zone group. For instance, Figure IM shows an Upper Area 109a including Zones A-D, and a Lower Area 109b including Zones E-I. In one aspect, an Area may be used to invoke a cluster of zone groups and / or zones that share one or more zones and / or zone groups of another cluster. In another aspect, this differs from a zone group, which does not share a zone with another zone group. Further examples of techniques for implementing Areas may be found, for example, in U.S. Application No. 15 / 682,506 filed August 21, 2017 and titled “Room Association Based on Name,” and U.S. Patent No. 8,483,853 filed September 11, 2007, and titled “Controlling and manipulating groupings in a multi-zone media system.” Each of these applications is incorporated herein by reference in its entirety. In some example configurations, the media playback system 100 may not implement Areas, in which case the system may not store variables associated with Areas.III. Example Systems and Devices

[0212] Figure 2A is a front isometric view of a playback device 210 configured in accordance with aspects of the disclosed technology. Figure 2B is a front isometric view of the playback device 210 without a grille 216e. Figure 2C is an exploded view of the playback device 210. Referring to Figures 2A-2C together, the playback device 210 comprises a housing 216 that includes an upper portion 216a, a right or first side portion 216b, a lower portion 216c, a left or second side portion 216d, the grille 216e, and a rear portion 216f. A plurality of fasteners 216g (e.g., one or more screws, rivets, clips) attaches a frame 216h to the housing 216. A cavity 216j (Figure 2C) in the housing 216 is configured to receive the frame 216h and electronics 212. The frame 216h is configured to carry a plurality of transducers 214 (identified individually in Figure 2B as transducers 214a-f). The electronics 212 (e.g., the electronics 112 of Figure 1C) is configured to receive audio content from an audio source and send electrical signals corresponding to the audio content to the transducers 214 for playback.

[0213] The transducers 214 are configured to receive the electrical signals from the electronics 112, and further configured to convert the received electrical signals into audible sound during playback. For instance, the transducers 214a-c (e.g., tweeters) can be configured tooutput high frequency sound (e.g., sound waves having a frequency greater than about 2 kHz). The transducers 214d-f (e.g., mid-woofers, woofers, midrange speakers) can be configured output sound at frequencies lower than the transducers 214a-c (e.g., sound waves having a frequency lower than about 2 kHz). In some example configurations, the playback device 210 includes a number of transducers different than those illustrated in Figures 2A-2C. For example, as described in further detail below with respect to Figures 3A-3C, the playback device 210 can include fewer than six transducers (e.g., one, two, three). In other example configurations, however, the playback device 210 includes more than six transducers (e.g., nine, ten). Moreover, in some example configurations, all or a portion of the transducers 214 are configured to operate as a phased array to desirably adjust (e.g., narrow or widen) a radiation pattern of the transducers 214, thereby altering a user’s perception of the sound emitted from the playback device 210.

[0214] In the illustrated embodiment of Figures 2A-2C, a filter 216i is axially aligned with the transducer 214b. The filter 216i can be configured to desirably attenuate a predetermined range of frequencies that the transducer 214b outputs to improve sound quality and a perceived sound stage output collectively by the transducers 214. In some example configurations, however, the playback device 210 omits the filter 216i . In other example configurations, the playback device 210 includes one or more additional filters aligned with the transducers 214b and / or at least another of the transducers 214.

[0215] Figures 3A and 3B are front and right isometric side views, respectively, of an NMD 320 configured in accordance with example configurations of the disclosed technology. Figure 3C is an exploded view of the NMD 320. Figure 3D is an enlarged view of a portion of Figure 3B including a user interface 313 of the NMD 320. Referring first to Figures 3A-3C, the NMD 320 includes a housing 316 comprising an upper portion 316a, a lower portion 316b and an intermediate portion 316c (e.g., a grille). A plurality of ports, holes or apertures 316d in the upper portion 316a allow sound to pass through to one or more microphones 315 (Figure 3C) positioned within the housing 316. The one or more microphones 316 are configured to received sound via the apertures 316d and produce electrical signals based on the received sound. In the illustrated embodiment, a frame 316e (Figure 3C) of the housing 316 surrounds cavities 316f and 316g configured to house, respectively, a first transducer 314a (e.g., a tweeter) and a second transducer 314b (e.g., a mid-woofer, a midrange speaker, a woofer). In other example configurations, however, the NMD 320 includes a single transducer, or more than two (e g., two,five, six) transducers. In certain example configurations, the NMD 320 omits the transducers 314a and 314b altogether.

[0216] Electronics 312 (Figure 3C) includes components configured to drive the transducers 314a and 314b, and further configured to analyze audio information corresponding to the electrical signals produced by the one or more microphones 315. In some example configurations, for example, the electronics 312 comprises many or all of the components of the electronics 112 described above with respect to Figure 1C. In certain example configurations, the electronics 312 includes components described above with respect to Figure IF such as, for example, the one or more processors 112a, the memory 112b, the software components 112c, the network interface 112d, etc. In some example configurations, the electronics 312 includes additional suitable components (e.g., proximity or other sensors).

[0217] Referring to Figure 3D, the user interface 313 includes a plurality of control surfaces (e.g., buttons, knobs, capacitive surfaces) including a first control surface 313a (e.g., a previous control), a second control surface 313b (e.g., a next control), and a third control surface 313c (e.g., a play and / or pause control). A fourth control surface 313d is configured to receive touch input corresponding to activation and deactivation of the one or microphones 315. A first indicator 313e (e.g., one or more light emitting diodes (LEDs) or another suitable illuminator) can be configured to illuminate only when the one or more microphones 315 are activated. A second indicator 313f (e.g., one or more LEDs) can be configured to remain solid during normal operation and to blink or otherwise change from solid to indicate a detection of voice activity. In some example configurations, the user interface 313 includes additional or fewer control surfaces and illuminators. In one embodiment, for example, the user interface 313 includes the first indicator 313e, omitting the second indicator 313f. Moreover, in certain example configurations, the NMD 320 comprises a playback device and a control device, and the user interface 313 comprises the user interface of the control device .

[0218] Referring to Figures 3A-3D together, the NMD 320 is configured to receive voice commands from one or more adjacent users via the one or more microphones 315. As described above with respect to Figure IB, the one or more microphones 315 can acquire, capture, or record sound in a vicinity (e.g., a region within 10m or less of the NMD 320) and transmit electrical signals corresponding to the recorded sound to the electronics 312. The electronics 312 can process the electrical signals and can analyze the resulting audio data to determine apresence of one or more voice commands (e.g., one or more activation words). Tn some example configurations, for example, after detection of one or more suitable voice commands, the NMD 320 is configured to transmit a portion of the recorded audio data to another device and / or a remote server (e.g., one or more of the computing devices 106 of Figure IB) for further analysis. The remote server can analyze the audio data, determine an appropriate action based on the voice command, and transmit a message to the NMD 320 to perform the appropriate action. For instance, a user may speak “Sonos, play Michael Jackson.” The NMD 320 can, via the one or more microphones 315, record the user’s voice utterance, determine the presence of a voice command, and transmit the audio data having the voice command to a remote server (e.g., one or more of the remote computing devices 106 of Figure IB, one or more servers of a VAS and / or another suitable service). The remote server can analyze the audio data and determine an action corresponding to the command. The remote server can then transmit a command to the NMD 320 to perform the determined action (e.g., play back audio content related to Michael Jackson). The NMD 320 can receive the command and play back the audio content related to Michael Jackson from a media content source. As described above with respect to Figure IB, suitable content sources can include a device or storage communicatively coupled to the NMD 320 via a LAN (e.g., the network 104 of Figure IB), a remote server (e.g., one or more of the remote computing devices 106 of Figure IB), etc. In certain example configurations, however, the NMD 320 determines and / or performs one or more actions corresponding to the one or more voice commands without intervention or involvement of an external device, computer, or server.

[0219] Figure 3E is a functional block diagram showing additional features of the NMD 320 in accordance with aspects of the disclosure. The NMD 320 includes components configured to facilitate voice command capture including voice activity detector component(s) 312k, beam forming components 3121, acoustic echo cancellation (AEC) and / or self-sound suppression components 312m, activation word detector components 312n, and voice / speech conversion components 312o (e.g., voice-to-text and text-to-voice). In the illustrated embodiment of Figure 3E, the foregoing components 312k-312o are shown as separate components. In some example configurations, however, one or more of the components 312k-312o are subcomponents of the processors 112a.

[0220] The beamforming and self-sound suppression components 3121 and 312m are configured to detect an audio signal and determine aspects of voice input represented in thedetected audio signal, such as the direction, amplitude, frequency spectrum, etc. The voice activity detector activity components 312k are operably coupled with the beamforming and AEC components 3121 and 312m and are configured to determine a direction and / or directions from which voice activity is likely to have occurred in the detected audio signal. Potential speech directions can be identified by monitoring metrics which distinguish speech from other sounds. Such metrics can include, for example, energy within the speech band relative to background noise and entropy within the speech band, which is measure of spectral structure. As those of ordinary skill in the art will appreciate, speech typically has a lower entropy than most common background noise.

[0221] The activation word detector components 312n are configured to monitor and analyze received audio to determine if any activation words (e.g., wake words) are present in the received audio. The activation word detector components 312n may analyze the received audio using an activation word detection algorithm. If the activation word detector 312n detects an activation word, the NMD 320 may process voice input contained in the received audio.Example activation word detection algorithms accept audio as input and provide an indication of whether an activation word is present in the audio. Many first- and third-party activation word detection algorithms are known and commercially available. For instance, operators of a voice service may make their algorithm available for use in third-party devices. Alternatively, an algorithm may be trained to detect certain activation words. In some example configurations, the activation word detector 312n runs multiple activation word detection algorithms on the received audio simultaneously (or substantially simultaneously). As noted above, different voice services (e g. AMAZON'S ALEXA®, APPLE'S SIRI®, or MICROSOFT'S CORTANA®) can each use a different activation word for invoking their respective voice service. To support multiple services, the activation word detector 312n may run the received audio through the activation word detection algorithm for each supported voice service in parallel.

[0222] The speech / text conversion components 312o may facilitate processing by converting speech in the voice input to text. In some example configurations, the electronics 312 can include voice recognition software that is trained to a particular user or a particular set of users associated with a household. Such voice recognition software may implement voiceprocessing algorithms that are tuned to specific voice profile(s). Tuning to specific voice profiles may require less computationally intensive algorithms than traditional voice activity services,which typically sample from a broad base of users and diverse requests that are not targeted to media playback systems.

[0223] Figure 3F is a schematic diagram of an example voice input 328 captured by the NMD 320 in accordance with aspects of the disclosure. The voice input 328 can include a activation word portion 328a and a voice utterance portion 328b. In some example configurations, the activation word 557a can be a known activation word, such as “Alexa,” which is associated with AMAZON'S ALEXA®. In other example configurations, however, the voice input 328 may not include a activation word. In some example configurations, a network microphone device may output an audible and / or visible response upon detection of the activation word portion 328a. In addition or alternately, an NMD may output an audible and / or visible response after processing a voice input and / or a series of voice inputs.

[0224] The voice utterance portion 328b may include, for example, one or more spoken commands (identified individually as a first command 328c and a second command 328e) and one or more spoken keywords (identified individually as a first keyword 328d and a second keyword 328f). In one example, the first command 328c can be a command to play music, such as a specific song, album, playlist, etc. In this example, the keywords may be one or words identifying one or more zones in which the music is to be played, such as the Living Room and the Dining Room shown in Figure 1A. In some examples, the voice utterance portion 328b can include other information, such as detected pauses (e.g., periods of non-speech) between words spoken by a user, as shown in Figure 3F. The pauses may demarcate the locations of separate commands, keywords, or other information spoke by the user within the voice utterance portion 328b.

[0225] In some example configurations, the media playback system 100 is configured to temporarily reduce the volume of audio content that it is playing while detecting the activation word portion 557a. The media playback system 100 may restore the volume after processing the voice input 328, as shown in Figure 3F. Such a process can be referred to as ducking, examples of which are disclosed in U.S. Patent Application No. 15 / 438,749, incorporated by reference herein in its entirety.

[0226] Figures 4A-4D are schematic diagrams of a control device 430 (e.g., the control device 130a of Figure 1H, a smartphone, a tablet, a dedicated control device, an loT device, and / or another suitable device) showing corresponding user interface displays in various states ofoperation. A first user interface display 431a (Figure 4A) includes a display name 433a (i.e., “Rooms”). A selected group region 433b displays audio content information (e.g., artist name, track name, album art) of audio content played back in the selected group and / or zone. Group regions 433c and 433d display corresponding group and / or zone name, and audio content information audio content played back or next in a playback queue of the respective group or zone. An audio content region 433e includes information related to audio content in the selected group and / or zone (i.e., the group and / or zone indicated in the selected group region 433b). A lower display region 433f is configured to receive touch input to display one or more other user interface displays. For example, if a user selects “Browse” in the lower display region 433f, the control device 430 can be configured to output a second user interface display 43 lb (Figure 4B) comprising a plurality of music services 433g (e.g., Spotify, Radio by Tunein, Apple Music, Pandora, Amazon, TV, local music, line-in) through which the user can browse and from which the user can select media content for play back via one or more playback devices (e.g., one of the playback devices 110 of Figure 1 A). Alternatively, if the user selects “My Sonos” in the lower display region 433f, the control device 430 can be configured to output a third user interface display 431c (Figure 4C). A first media content region 433h can include graphical representations (e.g., album art) corresponding to individual albums, stations, or playlists. A second media content region 433i can include graphical representations (e.g., album art) corresponding to individual songs, tracks, or other media content. If the user selections a graphical representation 433j (Figure 4C), the control device 430 can be configured to begin play back of audio content corresponding to the graphical representation 433j and output a fourth user interface display 43 Id fourth user interface display 43 Id includes an enlarged version of the graphical representation 433j, media content information 433k (e.g., track name, artist, album), transport controls 433m (e.g., play, previous, next, pause, volume), and indication 433n of the currently selected group and / or zone name.

[0227] Figure 5 is a schematic diagram of a control device 530 (e.g., a laptop computer, a desktop computer) . The control device 530 includes transducers 534, a microphone 535, and a camera 536. A user interface 531 includes a transport control region 533a, a playback status region 533b, a playback zone region 533c, a playback queue region 533d, and a media content source region 533e. The transport control region comprises one or more controls for controlling media playback including, for example, volume, previous, play / pause, next, repeat, shuffle, trackposition, crossfade, equalization, etc. The audio content source region 533e includes a listing of one or more media content sources from which a user can select media items for play back and / or adding to a playback queue.

[0228] The playback zone region 533b can include representations of playback zones within the media playback system 100 (Figures 1 A and IB). In some example configurations, the graphical representations of playback zones may be selectable to bring up additional selectable icons to manage or configure the playback zones in the media playback system, such as a creation of bonded zones, creation of zone groups, separation of zone groups, renaming of zone groups, etc. In the illustrated embodiment, a “group” icon is provided within each of the graphical representations of playback zones. The “group” icon provided within a graphical representation of a particular zone may be selectable to bring up options to select one or more other zones in the media playback system to be grouped with the particular zone. Once grouped, playback devices in the zones that have been grouped with the particular zone can be configured to play audio content in synchrony with the playback device(s) in the particular zone.Analogously, a “group” icon may be provided within a graphical representation of a zone group. In the illustrated embodiment, the “group” icon may be selectable to bring up options to deselect one or more zones in the zone group to be removed from the zone group. In some example configurations, the control device 530 includes other interactions and implementations for grouping and ungrouping zones via the user interface 531. In certain example configurations, the representations of playback zones in the playback zone region 533b can be dynamically updated as playback zone or zone group configurations are modified.

[0229] The playback status region 533c includes graphical representations of audio content that is presently being played, previously played, or scheduled to play next in the selected playback zone or zone group. The selected playback zone or zone group may be visually distinguished on the user interface, such as within the playback zone region 533b and / or the playback queue region 533d. The graphical representations may include track title, artist name, album name, album year, track length, and other relevant information that may be useful for the user to know when controlling the media playback system 100 via the user interface 531.

[0230] The playback queue region 533d includes graphical representations of audio content in a playback queue associated with the selected playback zone or zone group. In some example configurations, each playback zone or zone group may be associated with a playbackqueue containing information corresponding to zero or more audio items for playback by the playback zone or zone group. For instance, each audio item in the playback queue may comprise a uniform resource identifier (URI), a uniform resource locator (URL) or some other identifier that may be used by a playback device in the playback zone or zone group to find and / or retrieve the audio item from a local audio content source or a networked audio content source, possibly for playback by the playback device. In some example configurations, for example, a playlist can be added to a playback queue, in which information corresponding to each audio item in the playlist may be added to the playback queue. In some example configurations, audio items in a playback queue may be saved as a playlist. In certain example configurations, a playback queue may be empty, or populated but “not in use” when the playback zone or zone group is playing continuously streaming audio content, such as Internet radio that may continue to play until otherwise stopped, rather than discrete audio items that have playback durations. In some example configurations, a playback queue can include Internet radio and / or other streaming audio content items and be “in use” when the playback zone or zone group is playing those items.

[0231] When playback zones or zone groups are “grouped” or “ungrouped,” playback queues associated with the affected playback zones or zone groups may be cleared or reassociated. For example, if a first playback zone including a first playback queue is grouped with a second playback zone including a second playback queue, the established zone group may have an associated playback queue that is initially empty, that contains audio items from the first playback queue (such as if the second playback zone was added to the first playback zone), that contains audio items from the second playback queue (such as if the first playback zone was added to the second playback zone), or a combination of audio items from both the first and second playback queues. Subsequently, if the established zone group is ungrouped, the resulting first playback zone may be re-associated with the previous first playback queue, or be associated with a new playback queue that is empty or contains audio items from the playback queue associated with the established zone group before the established zone group was ungrouped. Similarly, the resulting second playback zone may be re-associated with the previous second playback queue, or be associated with a new playback queue that is empty, or contains audio items from the playback queue associated with the established zone group before the established zone group was ungrouped.

[0232] Figure 6 is a message flow diagram illustrating data exchanges between devices of the media playback system 100 (Figures 1A-1M).

[0233] At step 650a, the media playback system 100 receives an indication of selected media content (e.g., one or more songs, albums, playlists, podcasts, videos, stations) via the control device 130a. The selected media content can comprise, for example, media items stored locally on or more devices (e.g., the audio source 105 of Figure 1C) connected to the media playback system and / or media items stored on one or more media service servers (one or more of the remote computing devices 106 of Figure IB). In response to receiving the indication of the selected media content, the control device 130a transmits a message 651a to the playback device 110a (Figures 1A-1C) to add the selected media content to a playback queue on the playback device 110a.

[0234] At step 650b, the playback device 110a receives the message 651a and adds the selected media content to the playback queue for play back.

[0235] At step 650c, the control device 130a receives input corresponding to a command to play back the selected media content. In response to receiving the input corresponding to the command to play back the selected media content, the control device 130a transmits a message 65 lb to the playback device 110a causing the playback device 110a to play back the selected media content. In response to receiving the message 651b, the playback device 110a transmits a message 651c to the first computing device 106a requesting the selected media content. The first computing device 106a, in response to receiving the message 651c, transmits a message 65 Id comprising data (e.g., audio data, video data, a URL, a URI) corresponding to the requested media content.

[0236] At step 650d, the playback device 110a receives the message 65 Id with the data corresponding to the requested media content and plays back the associated media content.

[0237] At step 650e, the playback device 110a optionally causes one or more other devices to play back the selected media content. In one example, the playback device 110a is one of a bonded zone of two or more players (Figure IM). The playback device 110a can receive the selected media content and transmit all or a portion of the media content to other devices in the bonded zone. In another example, the playback device 110a is a coordinator of a group and is configured to transmit and receive timing information from one or more other devices in the group. The other one or more devices in the group can receive the selected media content fromthe first computing device 106a, and begin playback of the selected media content in response to a message from the playback device 110a such that all of the devices in the group play back the selected media content in synchrony.IV. Example Area Zone Configurations

[0238] As mentioned earlier, Area Zone configurations can provide enhanced flexibility and scalability over existing types of grouped configurations. In some instances, a single Area Zone may correspond to a single room. In some scenarios, however, there may not necessarily be a one-to-one relationship between a room and an Area Zone. For example, a single room may contain two or more Area Zones. Alternatively, an Area Zone may include playback entities in several rooms. Still further, a single playback entity may participate in more than one Area Zones.

[0239] Although aspects of the calibration procedures disclosed herein are described in the context of Area Zone configurations, many aspects of the disclosed calibration procedures are equally applicable to other types of groupings of playback entities, and one or more (or all) aspects of the disclosed calibration procedures could be used with other groupings of playback entities beyond Area Zones.

[0240] Figures 7A-7E show different example networks of playback entities configured into one or more Area Zones within one or more rooms. In operation, each playback entity in Figures 7A-7E can be any of (i) a physical playback device, (ii) a physical Multi-Player Playback Device, or (iii) a logical player implemented via one or more Multi-Player Playback Devices. However, for illustration purposes, in some examples described with reference to Figures 7A-7E, some of the playback entities are shown and described as physical playback devices while other playback entities are shown and described as Multi-Player Playback Devices that implement one or more logical players. As shown and described in more detail with reference to Figure 8, a Multi-Player Playback Device is a new type of playback device that includes one or more processors, tangible, non-transitory computer readable media and a plurality (e.g., up to eight, ten, twelve, or even more) configurable audio outputs, where the new Multi-Player Playback Device is configurable to implement from one to several logical players, and wherein an individual. In operation, one or more (or all) aspects of the disclosed calibration procedures areequally applicable to other types of playback devices, and one or more (or all) aspects of the disclosed calibration procedures could be used with other types of playback devices.

[0241] Figure 7A shows a network 700 of playback entities 701-706 configured into a single Area Zone 750 within a single room 780 according to some example embodiments. A controller 795 is configured to control aspects of the setup and operation of the network 700, including setup and operation of Area Zone 750 and the playback entities 701-703 within the Area Zone 750. The playback entities in network 700 may be any type of playback entity disclosed herein, and the controller 795 may be any type of controller disclosed herein.

[0242] Area Zone 750 includes playback entities 701-706. Playback entity 701 is configured as the Area Zone primary for Area Zone 750 as indicated by the heavy dashed outline of playback entity 701. Each of the playback entities 701-706 is shown as a single physical playback device.

[0243] In some embodiments, calibrating the playback entities in Area Zone 750 for operation in room 780 includes one or more or all of the above-described calibration processes. In some embodiments, calibrating the playback entities in Area Zone 750 for operation in room 780 includes (i) boundary condition determination, (ii) spectral and / or spatial analysis and calibration, and (iii) low frequency analysis and correction.

[0244] For example, calibrating the playback entities in Area Zone 750 may include a first step of performing a boundary condition determination. In the example shown in Figure 7A, playback entity 703 and playback entity 706 are in the comer of the room. And playback entities 701, 702, 704, and 705 are positioned near walls. Therefore, in this example configuration, the boundary condition determination step may be performed for all of the playback entities.

[0245] As explained earlier, in some embodiments, determining the boundary condition for an individual playback entity includes, for the playback entity: (i) emitting a calibration signal via one or more speakers associated with the playback entity, and (ii) receiving one or more acoustic reflections of the calibration signal. The acoustic reflections may be obtained via one or more microphones integrated within the playback entity, one or more microphones associated with the playback entity, one or more microphones of (or associated with) one more controller devices, and / or any other suitable microphone(s) separate from the playback entity and the controller device that is suitably positioned to measure reflections sufficient for determininga self-responses for the playback entity. In some instances, it can be advantageous for the microphone(s) capturing the reflections to be positioned in close proximity to the speaker(s) emitting the calibration signal(s).

[0246] In some embodiments, sensors within or associated with the playback entities can additionally or alternatively be used to detect boundary conditions. In embodiments that include ultrasonic transducers, using the ultrasonic transducers (i.e., emitters and / or detectors) to detect boundary conditions includes, for a playback entity: (i) emitting an ultrasonic signal via one or more ultrasonic emitters associated with the playback entity, and (ii) receiving one or more reflections of the ultrasonic signal via one or more ultrasonic detectors associated with the playback entity. In some instances, it can be advantageous for the ultrasonic detector(s) capturing the reflections to be positioned in close proximity to the ultrasonic emitter(s) emitting the ultrasonic signal(s).

[0247] In some embodiments, the arrangement shown in Figure 7A is depicted on a graphical user interface that shows the arrangement of the playback entities (or at least the speakers associated with the playback entities) in relation to the walls, corners, ceiling, and / or other reflective surfaces within the room 780. In practice, a system installer or other system operator can decide to detect boundary conditions for less than all of the playback entities in the Area Zone 750. For example, some of the playback entities (or their associated speakers) may not be positioned near any reflective surfaces, and there may be nominal benefits (or perhaps no benefit) to performing the boundary condition determination function for the playback entities having speakers that are not positioned near any reflective surfaces within the room 780.

[0248] In some instances, to help an installer or other system operator confirm that the arrangement of playback entities depicted within the graphical user interface is consistent with the actual arrangement of playback entities installed within the room 780, some playback entity embodiments include LED (or similar visual indicator) that the installer or system operator can use to correlate actual playback entities with the playback entity icons depicted within the graphical user interface. For example, in some instances, the LED (or similar indicator) on the playback entity is configured to flash in a particular pattern when the icon corresponding to that playback entity within the graphical user interface has been selected.

[0249] After the boundary condition determination, calibrating the playback entities in Area Zone 750 for operation in room 780 includes conducting a spectral and / or spatial analysis.

[0250] Performing the spectral analysis enables the acoustic characteristics of the room 780 to be identified, e.g., whether the room 780 tends to amplify or attenuate certain audio frequencies more or less than others. Once the acoustic characteristics of the room have been identified, individual playback entities are configured with equalization settings that account for the acoustic characteristics of the room 780. For example, the equalization settings may increase the gain applied to certain acoustic frequencies that the room 780 tends to attenuate and / or decrease the gain applied to (or perhaps apply negative gain to) other acoustic frequencies that the room 780 tends to amplify. In operation, playing audio with the equalization settings determined from the spatial analysis enables the playback entities to accounts for (e.g., offset) the acoustic characteristics of room 780, thereby improving sound of the audio playback experienced by listeners within the room 780.

[0251] In some embodiments, the spectral analysis is global for the entire Area Zone 750. In such embodiments, the spectral analysis is performed for the entire Area Zone, and the same equalization settings are applied to all of the playback entities within the Area Zone.

[0252] In other embodiments, the spectral analysis is localized for different areas within the Area Zone 750. A localized (rather than global) spectral analysis may be advantageous in scenarios where certain areas of the room 780 may tend to attenuate and / or amplify certain frequencies more or less than other areas of the room 780.

[0253] For example, one part of the room 780 with thick carpeting on the floor may tend to attenuate certain higher frequencies than another part of the room 780 that may have a hard floor (e g., concrete, wood, etc.). Thus, for a consistent sound within the room 780, it can be advantageous to have (i) a first set of equalization settings for the playback entities having speakers positioned in the area of the room 780 affected by the thick carpet and (ii) a second set of equalization settings for the playback entities having speakers positioned in an area of the room 780 affected by the hard flooring. In operation, the gain applied to the audio frequencies attenuated by the thick carpeting may be higher in the first equalization settings (applied to the audio played by the playback entities having speakers positioned in the area of the room attenuated by the thick carpet) than in the second set of equalization settings (applied to the audio played by the playback entities having speakers positioned in the area of the room amplified by the hard flooring).

[0254] In embodiments that also include spatial analysis and calibration, the spatial analysis includes determining playback timing delays to direct the time-of-arrival of audio to one or more spatial locations within the room 780.

[0255] After conducting the spectral and / or spatial analysis and calibration, calibrating the playback entities in Area Zone 750 for operation in room 780 includes low frequency analysis and correction.

[0256] In some embodiments, performing the low frequency analysis and correction includes causing all of the playback entities 701-706 to emit a broadband audio signal (e.g., white noise, pink noise, brown noise, or other suitable broadband audio signal) from their associated speakers. In some instances, this broadband signal is different from (i) high frequency signals used in the boundary condition determination and (ii) the frequency-sweeping calibration signals used in the spectral and / or spatial analysis and calibration. In some embodiments, the broadband signal comprises one or more bursts of broadband noise. In some embodiments, the broadband signal comprises an audio signal that varies in amplitude (or loudness) over time. In some embodiments, the broadband signal comprises an audio signal that remains constant for some fixed duration of time.

[0257] Performing the low frequency analysis and correction also includes obtaining a recording of the playback of the broadband signal within the room 780. Based on the recording of the playback of the broadband signal within the room 780, a decay time (e.g., RT60, which is defined as the measure of the time after the sound source ceases that it takes for the sound pressure level to reduce by 60 dB) is determined to assess an overall bass level output within the room 780. The recording used for determining the decay may be obtained via one or more microphones integrated within the playback entities, a microphone of a controller 795, and / or any other suitable microphone separate from the playback entities and the controller 795.

[0258] After determining the overall bass level output by the playback entities of the Area Zone 750 within the room 780, whether and the extent to which the determined overall bass level differs from a target bass level can be determined. When the determined overall bass level differs from the target bass level by more than a threshold amount, the bass equalization levels of one or more playback entities can be adjusted to cause the overall bass level to more closely match by the target bass level.

[0259] Some embodiments may include performing the low frequency analysis and correction several times in an iterative manner. For example, the overall bass level can be determined and compared to the target bass level for the room. The bass equalization level(s) for one or more playback entities can then be adjusted, and the process can be executed again with the adjusted bass level. In particular, the playback entities play the broadband audio signal using the adjusted bass settings from the previous iteration, and the overall bass level is then determined and compared to the target bass level to determine whether further changes should be made to the bass equalization level(s) for one or more of the playback entities. And if further changes to the bass equalization level(s) for one or more of the playback entities should be made, then the those changes are made and the process can repeat again. In some embodiments, the low frequency analysis and correction process is executed in this iterative manner until the overall measured bass level is sufficiently close to the target bass level for the room.

[0260] In some instances, the overall bass level may be different in different parts of the room 780 because of several factors, e.g., the dimensions of the room, the position(s) of the speakers of the playback entities, or other factors that can affect where and how bass frequencies tend to accumulate within a listening area.

[0261] For example, some embodiments include performing several localized low frequency analyses of different areas within a room, and then customizing equalizations for the different playback entities within the different areas of the room perhaps on an area-by-area basis. In this manner, such embodiments include performing multiple localized low frequency analyses for a single room, and then tailoring the bass equalization levels for individual playback entities in the different areas on an area-by-area basis within the room 780.

[0262] Figure 7B shows a network 720 of playback entities 701-712 configured into a two Area Zones 750-751 within room 781. Room 781 in Figure 7B is larger than room 780 in Figure 7A.

[0263] Area Zone 750 in Figure 7B is the same Area Zone 750 shown in Figure 7A in that it includes playback entities 701-706. Playback entity 701 is configured as the Area Zone primary for Area Zone 750 as indicated by the heavy dashed outline of playback entity 701. Each of the playback entities 701-706 is shown as a single physical playback device.

[0264] Area Zone 751 includes playback entities 707-712. Playback entity 707 is configured as the Area Zone primary for Area Zone 751 as indicated by the heavy dashed outlineof playback entity 707. Each of the playback entities 707-712 is shown as a single physical playback device.

[0265] In some embodiments, calibrating the playback entities in Area Zones 750 and 751 for operation in room 781 includes one or more or all of the above-described calibration processes. In some embodiments, calibrating the playback entities in Area Zones 750 and 751 for operation in room 781 includes (i) boundary condition determination, (ii) spectral and / or spatial analysis and calibration, and (iii) low frequency analysis and correction.

[0266] In some embodiments, the playback entities in Area Zone 750 are calibrated separately from the playback entities in Area Zone 751. In other embodiments, the boundary condition determination function and the spectral analysis and calibration function are performed first for Area Zone 750, and then second for Area Zone 751 separately from Area Zone 750. And then the low frequency analysis and correction is performed for Area Zone 750 and Area Zone 751 together (at the same time) since all of the playback entities in Area Zone 750 and Area Zone 751 are in the same room 781, and thus, the low frequencies of the audio played by the playback entities in Area Zone 750 can affect the audio experienced by listeners to the audio played by Area Zone 751, and vice versa.

[0267] For example, in some embodiments, calibrating the playback entities for Area Zone 750 and Area Zone 751 includes (i) determining boundary conditions for the playback entities in Area Zone 750, (ii) determining boundary conditions for the playback entities in Area Zone 751, (iii) conducting the spectral and / or spatial analysis and calibration for the playback entities of Area Zone 750, (iv) conducting the spectral and / or spatial analysis and calibration for the playback entities of Area Zone 751, and (iv) conducting the low frequency analysis and correction for all of the playback entities together in Area Zones 750 and 751 at the same time.

[0268] In some embodiments, determining the boundary conditions for the playback entities in Area Zone 750 may include determining boundary conditions for playback entities 701, 702, 703, and 706 since they are near walls, but foregoing determining boundary conditions for playback entities 704 and 705 since they are not near walls. However, in other embodiments, determining the boundary conditions for the playback entities in Area Zone 750 may include determining boundary conditions for all of the playback entities 701-706 in Area Zone 750 regardless of whether and the extent to which the individual playback entities 701-706 are near walls or not.

[0269] Similarly, in some embodiments, determining the boundary conditions for the playback entities in Area Zone 751 may include determining boundary conditions for playback entities 709, 710, 711, and 712 since they are near walls, but foregoing determining boundary conditions for playback entities 707 and 708 since they are not near walls. However, in other embodiments, determining the boundary conditions for the playback entities in Area Zone 751 may include determining boundary conditions for all of the playback entities 707-712 in Area Zone 751 regardless of whether and the extent to which the individual playback entities 707-712 are near walls or not.

[0270] Regardless of which playback entities for which the boundary condition determination is performed (i.e., all of the playback entities or fewer than all of the playback entities), the boundary condition determination for each playback entities is the same as any of the boundary condition determination processes described herein, including any alternative embodiments thereof. In particular, for an individual playback entity, the boundary condition determination is be performed by one or both of (i) playing acoustic calibration signals via the speaker(s) and recording the acoustic reflections for analysis and / or (ii) emitting ultrasonic signals and measuring the reflections of the ultrasonic signals for analysis. Then, if one or more aspects of the reflected signals (acoustic and / or ultrasonic) differ from one or more expected aspects of the reflected signals by more than a threshold amount, one or more equalization parameters of the playback entity are adjusted to correct for the difference between the actual aspects of the reflected signals and the expected aspects of the reflected signals. In this manner, the equalization settings for audio playback are adjusted to avoid or at least somewhat ameliorate the effect of acoustic reflections.

[0271] After determining boundary conditions for the playback entities in Area Zone 750 and Area Zone 751, the process of calibrating the playback entities next proceeds to conducting the spectral and / or spatial analysis and calibration for the playback entities of Area Zone 750 following by conducting the spectral and / or spatial analysis and calibration for the playback entities of Area Zone 751.

[0272] In operation, the spectral analysis and calibration may be performed for Area Zone 750 on a global level. Alternatively, the spectral analysis and calibration may be performed in the different regions of Area Zone 750 separately similar to the manner described above with reference to Figure 7A. Similarly, the spectral analysis and calibration may likewisebe performed for Area Zone 751 on a global level. Alternatively, the spectral analysis and calibration may be performed in the different regions of Area Zone 751 separately similar to the manner described above with reference to Figure 7A.

[0273] Some embodiments may additionally include conducting a spatial analysis and calibration to direct the time-of-arrival of audio to one or more spatial locations within the room 781. For example, some embodiments include conducting a first spatial analysis and calibration to direct the time-of-arrival of audio played via the playback entities of Area Zone 750 to one or more spatial locations within room 781, and then conducting a second spatial analysis and calibration to direct the time-of-arrival of audio played via the playback entities of Area Zone 751 to one or more different spatial locations within room 781.

[0274] After conducting the spectral and / or spatial analysis and calibration, some embodiments next include performing a low frequency analysis and correction procedure. Some embodiments may include performing a first low frequency analysis and correction procedure for the playback entities of Area Zone 750 followed by a second low frequency analysis and correction procedure for the playback entities of Area Zone 751.

[0275] However, as mentioned above, some embodiments include performing the low frequency analysis and correction procedure on all of the playback entities within both Area Zone 750 and Area Zone 751 together.

[0276] In embodiments that include performing the low frequency analysis and correction procedure on all of the playback entities within both Area Zone 750 and Area Zone 751 together, performing the low frequency analysis and correction procedure includes first causing all of the playback entities 701-712 in Area Zone 750 and Area Zone 751 to play an audio signal from their speakers. In some embodiments, the audio signal comprises a broadband audio signal such as white noise, pink noise, brown noise, or any other suitable broadband audio signal. In some embodiments the audio signal comprises one or more bursts of broadband noise.

[0277] At least while the playback entities are playing the audio signal (e.g., the broadband audio signal) and for at least some amount of time afterwards, one or more microphones record the playback of the audio signal (e.g., the broadband audio signal) within the room 781 for analysis. In some embodiments, the recording of the audio playback is analyzed to determine a decay time. For example, in some instances, an RT60 value is determined, which is defined as the measure of the time after the sound source ceases that it takes for the soundpressure level to reduce by 60 dB. This RT60 measurement is used to assess an overall bass level output within the listening area. The audio for determining the decay may be obtained via one or more microphones integrated within the playback entities 701-712, a microphone of a controller 795, and / or any other suitable microphone separate from the playback entities 701-712 and the controller 795.

[0278] After measuring the decay time, one or more equalization settings (e.g., a bass level) can be adjusted on one or more (or all) playback entities 701-712 of Area Zone 750 and Area Zone 751 within room 781. In scenarios where the decay time is measured at different locations within room 781, equalization settings can be tailored to each of the different locations based on the measured decay time for different locations. In operation, tailoring the equalization settings to the different locations may include using different bass level settings for the different playback entities 701-712 at the different locations within the room 781 so as to avoid excessive bass build up within the room 781 in general, and within different locations within the room 781 in particular

[0279] Figure 7C shows a network 740 of playback entities 701-718 configured into two Area Zones 750 and 752 within rooms 782 and 783.

[0280] Area Zone 750 in Figure 7C is the same Area Zone 750 shown in Figures 7A and 7B and includes playback entities 701-706. Playback entity 701 is configured as the Area Zone primary for Area Zone 750 as indicated by the heavy dashed outline of playback entity 701. Each of the playback entities 701-706 is shown as a single physical playback device.

[0281] Area Zone 752 includes playback entities 707-718. Playback entity 707 is configured as the Area Zone primary for Area Zone 752 as indicated by the heavy dashed outline of playback entity 707. Each of the playback entities 707-718 is shown as a single physical playback device. In contrast to Area Zone 751 in Figure 7B, Area Zone 752 in Figure 7C includes playback entities located in two separate rooms: room 782 and room 783. In particular, Area Zone 752 includes playback entities 707-712 in room 782 and playback entities 713-718 in room 783.

[0282] In some embodiments, calibrating the playback entities in Area Zones 750 and 752 for operation in rooms 782 and 783 includes one or more or all of the above-described calibration processes. In some embodiments, calibrating the playback entities in Area Zones 750and 752 for operation in rooms 782 and 783 includes (i) boundary condition determination, (ii) spectral and / or spatial analysis and calibration, and (iii) low frequency analysis and correction.

[0283] In some embodiments, the playback entities 701-706 in Area Zone 750 are calibrated separately from the playback entities 707-718 in Area Zone 752. And in some embodiments, the playback entities of Area Zone 752 in room 782 (i.e., playback entities 707- 712) are calibrated separately from the playback entities of Area Zone 752 in room 783 (i.e., playback entities 713-718).

[0284] In other embodiments, calibrating the playback entities in Area Zones 750 and 752 for operation in rooms 782 and 783 includes (i) determining boundary conditions for the playback entities in Area Zone 750 (i.e., playback entities 701-716), (ii) determining boundary conditions for the playback entities in Area Zone 752 that are in room 782 (i.e., playback entities 707-712), (iii) determining boundary conditions for the playback entities in Area Zone 752 that are in room 783 (i.e., playback entities 713-718), (iv) conducting the spectral and / or spatial analysis and calibration for the playback entities of Area Zone 750, (v) conducting the spectral and / or spatial analysis and calibration for the playback entities in Area Zone 752 that are in room 782 (i.e., playback entities 707-712), (vi) conducting the spectral and / or spatial analysis and calibration for the playback entities in Area Zone 752 that are in room 783 (i.e., playback entities 713-718), (vii) conducting the low frequency analysis and correction for all of the playback entities in room 782 together at the same time (i.e., playback entities 701-706 in Area Zone 750 and playback entities 707-712 in Area Zone 752); and (viii) conducting the low frequency analysis and correction for all of the playback entities in room 783 (i.e., playback entities 713- 718 in Area Zone 752).

[0285] In some instances, rather than conducting a spectral analysis and calibration of the playback entities in Area Zone 750 separately from a spectral analysis and calibration of the playback entities in Area Zone 752 that are in room 781, some embodiments include conducting a single spectral analysis and calibration for all of the playback entities configured to play audio in room 782 (i.e., a common spectral analysis and calibration for playback entities 701-706 in Area Zone 750 and playback entities 707-712 in Area Zone 752.

[0286] In operation, the boundary condition determination, spectral and / or spatial analysis and calibration, and low frequency analysis and correction for the individual playback entities in Figure 7C can be implemented according to any of the boundary conditiondetermination, spectral and / or spatial analysis and calibration, and low frequency analysis and correction disclosed and described herein.

[0287] Figure 7D shows a network of playback entities configured into two Area Zones 750 and 753 in two rooms 784 and 785.

[0288] Area Zone 750 in Figure 7D is the same Area Zone 750 shown in Figures 7A-7C and includes playback entities 701-706. Playback entity 701 is configured as the Area Zone primary for Area Zone 750 as indicated by the heavy dashed outline of playback entity 701. Each of the playback entities 701-706 is shown as a single physical playback device.

[0289] Area Zone 753 includes playback entities 707-709, 716-718, and 730-1 through 730-6 configured to play audio in room 784 and room 785. Each of the playback entities 707- 709 and 716-718 are single physical playback devices, where playback entities 707-709 are configured to play audio in room 784 and playback entities 716-718 are configured to play audio in room 785.

[0290] Area Zone 753 also includes Multi-Player Playback Device 730 that includes logical players 730-1 through 730-6. Logical players 730-1, 730-2, and 730-3 are configured to play audio in room 784 and logical players 730-4, 730-5, and 730-6 are configured to play audio in room 785.

[0291] As described in detail with reference to Figure 8, a Multi-Player Playback Device is a type of physical playback device comprising multiple configurable audio outputs. One or more audio outputs of an Multi-Player Playback Device can be configured to operate as a logical player. In some embodiments, a Multi-Player Playback Device can implement from one to eight (or perhaps more) logical players. In the example shown in Figure 7D, logical players 730-1, 730-2, and 730-2 are configured to drive 3 speakers that are located in room 784, and logical players 730-4, 730-5, and 730-6 are configured to drive 3 speakers that are located in room 785. Thus, in this manner, from a listener’s perspective, logical players 730-1, 730-2, and 730-3 in Figure 7D are functionally similar to playback entities 710, 711, and 712 in Figure 7C, and logical players 730-4, 730-5, and 730-6 in Figure 7D are functionally similar to playback entities 713, 714, and 715 in Figure 7C.

[0292] In some embodiments, the playback entities 701-706 in Area Zone 750 are calibrated separately from the playback entities in Area Zone 753. In some embodiments, the playback entities of Area Zone 753 in room 784 (i.e., playback entities 707-709 and logicalplayers 730-1 through 730-3) are calibrated separately from the playback entities of Area Zone 753 in room 785 (i.e., playback entities 716-718 and logical players 730-4 through 730-6).

[0293] In some embodiments, calibrating the playback entities in Area Zones 750 and 753 for operation in rooms 784 and 785 includes (i) determining boundary conditions for the playback entities in Area Zone 750 (i.e., playback entities 701-716), (ii) determining boundary conditions for the playback entities in Area Zone 753 that are in room 784 (i.e., playback entities 707-709 and logical players 730-1 through 730-3), (iii) determining boundary conditions for the playback entities in Area Zone 753 that are in room 785 (i.e., playback entities 716-718 and logical players 730-4 through 730-6), (iv) conducting the spectral and / or spatial analysis and calibration for the playback entities of Area Zone 750, (v) conducting the spectral and / or spatial analysis and calibration for the playback entities in Area Zone 753 that are in room 784 (i.e., playback entities 707-709 and logical players 730-1 through 730-3), (vi) conducting the spectral and / or spatial analysis and calibration for the playback entities in Area Zone 753 that are in room 785 (i.e., playback entities 716-718 and logical players 730-4 through 730-6), (vii) conducting the low frequency analysis and correction for all of the playback entities in room 784 together at the same time (i.e., playback entities 701-706 in Area Zone 750, playback entities 707-709, and logical players 730-1 through 730-3 in Area Zone 752); and (viii) conducting the low frequency analysis and correction for all of the playback entities in room 785 (i.e., playback entities 716- 718 and logical players 730-4 through 7306 in Area Zone 752).

[0294] In some instances, rather than conducting a spectral analysis and calibration of the playback entities in Area Zone 750 separately from a spectral analysis and calibration of the playback entities in Area Zone 753 that are in room 784, some embodiments include conducting a single spectral analysis and calibration for all of the playback entities configured to play audio in room 784 (i.e., a common spectral analysis and calibration for playback entities 701-706 in Area Zone 750, playback entities 707-709 in Area Zone 753, and logical players 730-1 through 730-3 in Area Zone 753.

[0295] In operation, the boundary condition determination, spectral and / or spatial analysis and calibration, and low frequency analysis and correction for the individual playback entities in Figure 7D can be implemented according to any of the boundary condition determination, spectral and / or spatial analysis and calibration, and low frequency analysis and correction disclosed and described herein.

[0296] Figure 7E shows a network 770 of playback entities configured into two Area Zones (i.e., Area Zone 754 and Area Zone 755) in two rooms (i.e., room 786 and room 787).

[0297] Area Zone 754 in room 786 includes (i) logical players 732-1 through 732-6 implemented by Multi-Player Playback Device 732 and (ii) logical players 730-1 through 730-6 implemented by Multi-Player Playback Device 730. Logical player 730-1 is configured as the Area Zone primary for Area Zone 754 as indicated by the heavy dashed outline of logical player 730-1.

[0298] Area Zone 755 in room 787 includes logical players 730-7 through 730-12 implemented by Multi-Player Playback Device 730. Logical player 730-7 is configured as the Area Zone primary for Area Zone 755 as indicated by the heavy dashed outline of logical player 730-7.

[0299] In the example shown in Figure 7E, logical players 732-1 through 732-6 are configured to drive 6 speakers that are located in room 786, and logical players 730-1 through 730-6 are configured to drive 6 additional speakers that are located in room 786. Thus, in this manner, from a listener’s perspective, logical players 732-1 through 732-6 in Figure 7E are functionally similar to playback entities 701-706 in Figures 7A-7D, and logical players 730-1 through 730-6 in Figure 7E are functionally similar to playback entities 707-712 in Figure 7C.

[0300] Further, in the example shown in Figure 7E, logical players 732-7 through 730-12 are configured to drive 6 speakers that are located in room 787. Thus, from a listener’s perspective, logical players 730-7 through 730-12 are functionally similar to playback entities 713-718 in Figure C.

[0301] In some embodiments, the playback entities Area Zone 754 (i.e., logical players 732-1 through 732-6 and logical players 730-1 through 730-6) are calibrated separately from the playback entities in Area Zone 755 (i.e., logical players 730-7 through 730-12).

[0302] In some embodiments, calibrating the playback entities in Area Zones 754 and 755 for operation in rooms 786 and 787 includes (i) determining boundary conditions for the playback entities in Area Zone 754 (i.e., logical players 732-1 through 732-6 and logical players 730-1 through 730-6), (ii) determining boundary conditions for the playback entities in Area Zone 755 (i.e., logical players 730-7 through 730-12), (iii) conducting the spectral and / or spatial analysis and calibration for the playback entities of Area Zone 754 (i.e., logical players 732-1 through 732-6 and logical players 730-1 through 730-6), (iv) conducting the spectral and / orspatial analysis and calibration for the playback entities in Area Zone 755 (i.e., logical players 730-7 through 730-12), (v) conducting the low frequency analysis and correction for all of the playback entities in room 786 together at the same time (i.e., logical players 732-1 through 732-6 and logical players 730-1 through 730-6); and (vi) conducting the low frequency analysis and correction for all of the playback entities in room 787 together at the same time (i.e., logical players 730-7 through 730-12).

[0303] In operation, the boundary condition determination, spectral and / or spatial analysis and calibration, and low frequency analysis and correction for the individual playback entities in Figure 7E can be implemented according to any of the boundary condition determination, spectral and / or spatial analysis and calibration, and low frequency analysis and correction disclosed and described herein.V. Example Multi-Player Playback Device

[0304] Figure 8 shows an example block diagram of a Multi-Player Playback Device 800 configurable to implement from one to eight logical players according to some example embodiments. The Multi-Player Playback Devices shown in Figures 7D-E may be the same as or similar to Multi-Player Playback Device 800 shown in Figure 8.

[0305] As mentioned earlier, a Multi-Player Playback Device is a playback device comprising multiple configurable audio outputs, one or more processors, and tangible, non- transitory computer readable media storing program instructions that are executed by the one or more processors to cause the Multi-Player Playback Device to perform the Multi-Player Playback Device features and functions described herein.

[0306] For example, Multi-Player Playback Device 800 includes one or more processors 802 and one or more tangible, non-transitory computer-readable memory 804. The tangible, non-transitory computer-readable memory 804 is configured to store program instructions that are executable by the one or more processors 802. When executed by the one or more processors 802, the program instructions cause the Multi-Player Playback Device 800 to perform the Multi- Player Playback Device functions disclosed and described herein.

[0307] Multi-Player Playback Device 800 also includes one or more network interfaces 806. The network interfaces may include any one or more (i) wired network interfaces, e.g., Ethernet, Universal Serial Bus, Firewire, Power-over-Ethernet, or any other type of wirednetwork interface now known or later developed that is suitable for transmitting and receiving the types of data described herein and / or (ii) wireless network interfaces, e.g., WiFi, Bluetooth, 4G / 5G, or any other type of wireless interface now known or later developed that is suitable for transmitting and receive the types of data described herein.

[0308] Multi-Player Playback Device includes amplifiers 810-1, 810-2, 810-3, and 810- 4, where each amplifier is connected to two audio outputs. In particular, (i) amplifier 810-1 is connected to audio outputs 812-1 and 812-2, (ii) amplifier 810-2 is connected to audio outputs 812-3 and 812-4, (iii) amplifier 810-3 is connected to audio outputs 812-5 and 812-6; and (iv) amplifier 810-4 is connected to audio outputs 812-7 and 812-8. Multi-Player Playback Device 800 includes four amplifiers 810-1 through 810-4 and eight audio outputs 812-1 through 812-8 as an illustrative example. Other embodiments may have a different number of amplifiers, a different number of audio outputs, and / or a different ratio of amplifiers to audio outputs.

[0309] In operation, the Multi-Player Playback Device 800 is configurable to implement from one to eight “logical” players. Each logical player (sometimes referred to herein as simply a player) can be addressed and managed within a playback system as a distinct playback entity.

[0310] In some embodiments, the Multi-Player Playback Device 800 is configured to operate within any of several different operating modes. In some embodiments, the Multi-Player Playback Device 800 is reconfigurable to switch between operating in one mode to operating in a different mode.

[0311] For example, in some embodiments, the Multi-Player Playback Device 800 is configurable to operate in a first mode where the Multi-Player Playback Device 800 is configured to implement eight single-channel logical players, where each logical player is configured to output single-channel audio from one of the eight audio outputs.

[0312] In this first mode of operation, each of the audio outputs 812-1 through 812-8 plays the same channel of audio. In some embodiments, the single-channel audio comprises a mono audio stream. In some embodiments, the single-channel audio comprises one channel from a multi-channel audio stream. In some scenarios, each of the eight logical players can be connected to a different speaker (e.g., free-standing speakers, speakers mounted in walls or ceilings, and so on) so that each connected speaker plays the same channel of audio. In some embodiments, all eight of the logical players are configured to output the same audio stream. Inother embodiments, one or more of the logical players may be configured to output different audio than one or more other logical players.

[0313] In some embodiments, the Multi-Player Playback Device 800 is configurable to operate in a second mode where the Multi-Player Playback Device 800 is configured to implement four two-channel logical players, wherein each logical player is configured to output two-channel audio from two of the eight audio outputs.

[0314] In this second mode of operation, (i) a first logical player uses amp 810-1 to drive audio outputs 812-1 and 812-2 to output two channels of audio, (ii) a second logical player uses amp 810-2 to drive audio outputs 812-3 and 812-4 to output two channels of audio, (ii) a third logical player uses amp 810-3 to drive audio outputs 812-5 and 812-6 to output two channels of audio, and (iv) a fourth logical player uses amp 810-4 to drive audio outputs 812-7 and 812-8 to output two channels of audio. The two channels of audio could be left and right channels of a stereo audio stream. Alternatively, the two channels of audio could be any two different channels of a multichannel audio stream. In some embodiments, all four of the logical players are configured to output the same audio stream. In other embodiments, one or more of the logical players may be configured to output different audio than one or more other logical players.

[0315] In some embodiments, the Multi-Player Playback Device 800 is configurable to operate in a third mode where the Multi-Player Playback Device 800 is configured to implement two four-channel logical players, wherein each logical player is configured to output four- channel audio from four of the eight audio outputs.

[0316] In this third mode of operation, (i) a first logical player uses amps 810-1 and 810- 2 to drive audio outputs 812-1, 812-2, 812-3, and 812-4 to output four channels of audio, and (ii) a second logical player uses amps 810-3 and 810-4 to drive audio outputs 812-5, 812-6, 812-7, and 812-8 to output four channels of audio. The four channels of audio could be front left, front right, rear left, and rear right channels of a surround sound audio stream. Alternatively, the four channels of audio could be any four different channels of a multichannel audio stream. In some embodiments, both of the logical players are configured to output the same audio stream. In other embodiments, one of the logical players may be configured to output different audio than the other logical player.

[0317] In some embodiments, the Multi-Player Playback Device 800 is configurable to operate in a fourth mode where the Multi-Player Playback Device 800 is configured to implement one eight-channel logical player, where the logical player is configured to output eight-channel audio from the eight audio outputs.

[0318] In this fourth mode of operation, a single logical player uses amps 810-1 through 810-4 to drive audio outputs 812-1 through 812-8 to output eight channels of audio. The eight channels of audio could be front left, front center, front right, rear left, rear center, rear right, left subwoofer, and right subwoofer channels of a surround sound audio stream. Alternatively, the eight channels of audio could be any eight different channels of a multichannel audio stream. In some embodiments, the logical player is configured to output the same channel of audio via each of the eight audio outputs.

[0319] Although four operating modes are described here for illustrative purposes, the Multi-Player Playback Device 800 in some embodiments is configurable to operate in other operating modes that include combinations of differently-configured players. For example, in some embodiments, the Multi-Player Playback Device 800 is configurable to operate in a mode that includes (i) a first logical player that uses amps 810-1, 810-2, and 810-3 to drive audio outputs 812-1, 812-2, and 812-3 to output six channels of audio, and (ii) a second logical player uses amp 810-4 to drive audio output 812-8 to output one channel of audio. The Multi-Player Playback Device 800 can implement other combinations and configurations of amps and audio outputs to output different channels of audio content.

[0320] In some embodiments, several Multi-Player Playback Devices (e.g., several Multi-Player Playback Devices like Multi-Player Playback Device 800) can be bonded together to implement more sophisticated configurations of one or more logical players.

[0321] For example, in some embodiments, the Multi-Player Playback Device 800 is configured to selectively operate in any of a plurality of bonded modes comprising a first bonded mode and a second bonded mode.

[0322] In the first bonded mode, the Multi-Player Playback Device 800 is configured to operate in a bonded configuration with a second Multi-Player Playback Device that has eight audio outputs similar to Multi-Player Playback Device 800. In the first bonded mode, the Multi- Player Playback Device 800 and the second Multi-Player Playback Device are configured to implement a single logical player configured to output any of (i) one set of sixteen channels ofaudio, (ii) two sets of eight channels of audio, (iii) four sets of four channels of audio, (iv) eight sets of two channels of audio, or (v) sixteen single channels of audio.

[0323] In the second bonded mode, the Multi-Player Playback Device 800 operate in a bonded configuration with a second Multi-Player Playback Device and a third Multi-Player Playback Device, where the second Multi-Player Playback Device and the third Multi-Player Playback Device each have eight audio outputs similar to Multi-Player Playback Device 800. In the second bonded mode, the Multi-Player Playback Device 800, the second Multi-Player Playback Device, and the third Multi-Player Playback Device are configured to implement a single logical player configured to output any of (i) one set of twenty four channels of audio, (ii) two sets of twelve channels of audio, (iii) three sets of eight channels of audio, (iv) four sets of six channels of audio, (v) six sets of four channels of audio, (v) eight sets of three channels of audio, (vi) twelve sets of two channels of audio, or (vii) twenty four singles channels of audio.VI. Example Methods

[0324] Figures 9A-C shows steps of an example method 900 depicting aspects of calibrating playback entities in playback networks comprising multiple playback entities according to some example embodiments.

[0325] For illustration purposes, example method 900 is described in the context of embodiments where a computing device or system of computing devices is configured to cause one or more playback entities (including one or more physical playback devices, physical Multi- Player Playback Devices, and logical players that are implemented via one or more Multi-Player Playback Devices) in an Area Zone comprising a first playback entity and one or more other playback entities located in a listening environment. Also for illustration purposes, method 900 is described in the context of Area Zones. But aspects of method 900 and variations thereof are equally applicable to other types of groupings of playback entities.

[0326] However, one or more (or all) aspects of example method 900 could be performed by one playback entity individually or in combination with one or more other playback entities, computing devices, and / or computing systems, including, for example, computing devices and / or computing systems separate from the playback entity / entities in the playback network.

[0327] For example, one or more (or all) aspects of method 900 may be performed by any one or more (or all) of the playback entities described with reference to the exampleplayback networks shown and described with reference to Figures 7A-E, and 8. In another example, one or more (or all) aspects of method 900 may be performed by a computing device and / or computing system separate from the playback entities in a playback network that includes one or more Area Zones.

[0328] In some embodiments, method 900 is performed by a computing device. In some embodiments, the computing device comprises one of (i) a computing device separate from the first playback entity and the one or more other playback entities in the area zone configuration,(ii) a computing entity of a cloud-based computing system that is configured to control aspects of the operation of the area zone configuration, including causing playback entities to emit sounds, record sounds and provide the recorded sounds to the cloud-based computing for analysis, and transmit commands and / or configuration files to one the playback entities for implementation,(iii) the first playback entity, or (iv) one of the one or more other playback entities in the area zone configuration.

[0329] Method 900 begins at method block 902, which includes causing the first playback entity to emit a first audio signal via one or more speakers associated with the first playback entity. In some embodiments, the first playback entity may cause itself to emit the first audio signal. In some examples, the one or more speakers associated with the first playback entity are speakers integrated with the first playback entity. In other examples, the one or more speakers are separate from the first playback entity, but connected to the playback entity in a manner sufficient to enable the first playback entity to play audio via the speakers.

[0330] In some embodiments, method block 902 additionally or alternatively includes: (i) causing the first playback entity to emit an ultrasonic audio signal; (ii) determining whether one or more aspects of detected reflections of the ultrasonic signal differ from one or more expected aspects of the reflections by more than a threshold amount; and (iii) when one or more aspects of the detected reflections differ from one or more expected aspects of the reflections by more than a threshold amount, adjust one or more equalization parameters of the first playback entity.

[0331] Next, method 900 advances to method block 904, which includes obtaining a first recording from at least one microphone associated with the first playback entity, wherein the first recording comprises reflections of the first audio signal from one or more objects in the listening environment. In some embodiments, the first playback entity may cause itself to obtain the first recording from at least one microphone associated with the first playback entity. In someexamples, the at least one microphone associated with the first playback entity comprises at least one microphone integrated with the first playback entity. In other examples, the at least one microphone is separate from the first playback entity, but connected to the playback entity in a manner sufficient to enable the first playback entity to obtain the first recording. Some embodiments additionally or alternatively include sending the first recording to another playback entity and / or one or more computing devices or systems for analysis.

[0332] Method 900 next proceeds to method block 906, which includes determining a self-response of the first playback entity based on the first recording.

[0333] In some embodiments, determining a self-response of the first playback entity based on the first recording at method block 906 includes determining whether one or more aspects of the reflections differ from one or more expected aspects of the reflections by more than a threshold amount. Some embodiments also include, when one or more aspects of the reflections differ from one or more expected aspects of the reflections by more than a threshold amount, adjusting one or more equalization parameters of the first playback entity.

[0334] In some embodiments, determining whether one or more aspects of the reflections differ from one or more expected aspects of the reflections by more than a threshold amount includes determining that the first playback entity’s proximity to at least one surface in the listening environment caused the one or more aspects of the reflections to differ from the one or more expected aspects of the reflections by more than the threshold amount. In some embodiments, the at least one surface comprises one of a ceiling, a wall, or a floor within the listening area. Additionally, in some embodiments, when one or more aspects of the reflections differ from one or more expected aspects of the reflections by more than a threshold amount, adjusting one or more equalization parameters of the first playback entity includes adjusting one or more equalization parameters of the first playback entity based on the first playback entity’s proximity to the at least one surface in the listening environment.

[0335] Next, method 900 advances to method block 908, which includes causing the first playback entity to emit a second audio signal via the one or more speakers while each of the one or more other playback entities in the area zone configuration also emit the second audio signal, wherein the second audio signal is different than the first audio signal.

[0336] After block 908, method 900 proceeds to method block 910, which includes obtaining a second recording of the second audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration.

[0337] In some embodiments, obtaining the second recording of the second audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration at method block 910 includes causing the computing device to obtain the second recording from the at least one microphone associated with the first playback entity.

[0338] In other embodiments, obtaining the second recording of the second audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration at method block 910 includes causing the computing device to obtain the second recording from a microphone separate from the computing device.

[0339] Method 900 then proceeds to method block 912, which includes determining a bass level in the listening environment based on the second recording.

[0340] In some embodiments, the second audio signal of method block 908 comprises a broadband signal. In embodiments where the second audio signal of method block 908 comprises a broadband signal, the step of determining the bass level in the listening environment based on the second recording at method block 912 includes causing the computing device to determine the bass level in the listening environment based on an RT60 analysis of the second recording.

[0341] Next, method 900 advances to method block 914, which includes, when the determined bass level in the listening environment differs from a target bass level for the listening environment by more than a threshold amount, adjusting a bass equalization level of at least one playback entity in the area zone configuration.

[0342] Next, method 900 advances to method block 916 (Figure 9B), which includes causing the first playback entity to emit a third audio signal via the one or more speakers while each of the one or more other playback entities in the area zone configuration also emit the third audio signal, wherein the third audio signal is different than the first audio signal and the second audio signal.

[0343] Method 900 next advances to method block 918, which includes obtaining a third recording of the third audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration.

[0344] In some embodiments, obtaining the third recording of the third audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration at method block 918 includes causing the computing device to obtain the third recording from the at least one microphone associated with the first playback entity. Some embodiments may additionally or alternatively include causing the computing device to obtain the third recording from the at least one microphone of one or more other playback entities in the area zone configuration.

[0345] In some embodiments, obtaining the third recording of the third audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration at method block 918 includes cause the computing device to obtain the third recording from a microphone separate from the computing device.

[0346] In some embodiments, obtaining the third recording of the third audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration at method block 918 includes obtaining one or more signals from one or more microphones at different locations within the listening environment. In some embodiments, the one or more microphones at the different locations within the listening environment comprise at least one or more of (i) a microphone at the computing device, (ii) a microphone at the first playback entity, (iii) a microphone at one of the one or more other playback entities in the area zone configuration, or (iv) a microphone separate from the computing device or the playback entities in the area zone configuration.

[0347] Next, method 900 proceeds to method block 920, which includes determining an acoustic response of the listening environment based on the third recording.

[0348] Next, method 900 advances to method block 922, which includes when the determined acoustic response of the listening environment differs from a target acoustic response for the listening environment by more than a threshold amount, adjusting one or more equalization levels of at least one playback entity in the area zone configuration.

[0349] In some embodiments, method 900 additionally executes the functions of method blocks 924-928 shown in Figure 9C. In operation, the additional functions performed in method blocks 924-928 in Figure 9C amount to performing the same or similar functions of method blocks 902-914 of Figure 9A for one or more other playback entities (or perhaps all of the other playback entities) within the area zone configuration.

[0350] In such embodiments, method 900 advances to method block 924, which includes, for at least one additional playback entity of the one or more additional playback entities of the one or more other playback entities in the area zone configuration, and when each of the other playback entities in the area zone configuration is not emitting an audio signal: (i) causing the additional playback entity to emit a self-response audio signal via one or more speakers associated with the additional playback entity, (ii) obtaining a recording comprising reflections of the self-response audio signal from one or more objects in the listening environment, and (iii) determining one or more boundary conditions of the additional playback entity based on the recording comprising the reflections of the self-response audio signal.

[0351] In some embodiments, the step of obtaining a recording comprising reflections of the self-response audio signal from one or more objects in the listening environment at method block 924 includes causing the computing device to obtain a separate recording from at least one microphone at each playback entity in the area zone configuration, wherein the second recording comprises each of the separate recordings obtained from each of the playback entities in the area zone configuration.

[0352] Next, method 900 proceeds to method block 926, which includes determining whether one or more aspects of the reflections in the recording obtained in block 924 differ from one or more expected aspects of the reflections in the recording obtained in block 924 by more than a threshold amount.

[0353] Method 900 next advances to method block 928, which includes when one or more aspects of the reflections in the recording obtained in block 924 differ from one or more expected aspects of the reflections in the recording obtained in block 924 by more than a threshold amount, adjust one or more equalization parameters of the additional playback entity.

[0354] In some embodiments, the one or more additional playback entities of the one or more other playback entities in the area zone configuration referred to in method blocks 924-928 of Figure 9C comprise fewer than all of the playback entities in the area zone configuration.

[0355] Some embodiments of method 900 additionally include causing a graphical user interface (e.g., a graphical user interface associated with a computing device configured to control one or more operational aspects of the area zone configuration) to display a representation of the playback entities of the area zone configuration. In some embodiments, individual playback entities are represented by icons that are selectable via user inputs to thegraphical user interface. In operation, an individual playback entity can be selected via a user input to the graphical user interface, and after receiving one or more commands to calibrate the individual playback entity (or perhaps the area zone overall), the computing device associated with the graphical user interface causes one or more playback entities to perform aspects of the functions disclosed and described with reference to Figures 9A-C.VII. Other Example Embodiments

[0356] Example 1: A computing device comprising one or more processors, and tangible, non-transitory computer-readable media comprising program instructions executable by the one or more processors such that the computing device is configured to, for an area zone configuration comprising a first playback entity and one or more other playback entities located in a listening environment: (i) cause the first playback entity to emit a first audio signal via one or more speakers of the first playback entity; (ii) obtain a first recording from at least one microphone of the first playback entity, wherein the first recording comprises reflections of the first audio signal from one or more objects in the listening environment; (iii) determine a selfresponse of the first playback entity based on the first recording; (iv) cause the first playback entity to emit a second audio signal via the one or more speakers while each of the one or more other playback entities in the area zone configuration also emit the second audio signal, wherein the second audio signal is different than the first audio signal; (v) obtain a second recording of the second audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration; (vi) determine a bass level in the listening environment based on the second recording; and (vii) when the determined bass level in the listening environment differs from a target bass level for the listening environment by more than a threshold amount, adjust a bass equalization level of at least one playback entity in the area zone configuration.

[0357] Example 2: The computing device of example 1, wherein the program instructions comprise program instructions executable by the one or more processors such that the computing device is configured to: (i) cause the first playback entity to emit an ultrasonic audio signal; (ii) determine whether one or more aspects of detected reflections of the ultrasonic signal differ from one or more expected aspects of the reflections by more than a threshold amount; and (iii) when one or more aspects of the detected reflections differ from one or moreexpected aspects of the reflections by more than a threshold amount, adjust one or more equalization parameters of the first playback entity.

[0358] Example 3: The computing device of example 2, wherein: (i) the program instructions executable by the one or more processors such that the computing device is configured to determine whether one or more aspects of the reflections differ from one or more expected aspects of the reflections by more than a threshold amount comprise program instructions executable by the one or more processors such that the computing device is configured to cause the computing device to determine that the first playback entity’s proximity to at least one surface in the listening environment caused the one or more aspects of the reflections to differ from the one or more expected aspects of the reflections by more than the threshold amount, wherein the at least one surface comprises one of a ceiling, a wall, or floor; and (ii) the program instructions executable by the one or more processors such that the computing device is configured to when one or more aspects of the reflections differ from one or more expected aspects of the reflections by more than a threshold amount, adjust one or more equalization parameters of the first playback entity comprise program instructions executable by the one or more processors such that the computing device is configured to cause the computing device to adjust one or more equalization parameters of the first playback entity based on the first playback entity’s proximity to the at least one surface in the listening environment.

[0359] Example 4: The computing device of example 1, wherein the program instructions executable by the one or more processors such that the computing device is configured to obtain the second recording of the second audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration comprise program instructions executable by the one or more processors such that the computing device is configured to cause the computing device to obtain the second recording from the at least one microphone of the first playback entity.

[0360] Example 5: The computing device of example 4, wherein the program instructions executable by the one or more processors such that the computing device is configured to obtain the second recording of the second audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration comprise program instructions executable by the one or more processors such thatthe computing device is configured to cause the computing device to obtain the second recording from a microphone separate from the computing device.

[0361] Example 6: The computing device of example 1, wherein the second audio signal comprises a broadband signal, and wherein the program instructions executable by the one or more processors such that the computing device is configured to determine the bass level in the listening environment based on the second recording comprise program instructions executable by the one or more processors such that the computing device is configured to cause the computing device to determine the bass level in the listening environment based on an RT60 analysis of the second recording.

[0362] Example 7: The computing device of example 1, wherein the program instructions comprise program instructions executable by the one or more processors such that the computing device is configured to, for the area zone configuration comprising the first playback entity and one or more other playback entities located in the listening environment: (i) cause the first playback entity to emit a third audio signal via the one or more speakers while each of the one or more other playback entities in the area zone configuration also emit the third audio signal, wherein the third audio signal is different than the first audio signal and the second audio signal; (ii) obtain a third recording of the third audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration; (iii) determine an acoustic response of the listening environment based on the third recording; and (iv) when the determined acoustic response of the listening environment differs from a target acoustic response for the listening environment by more than a threshold amount, adjust one or more equalization levels of at least one playback entity in the area zone configuration.

[0363] Example 8: The computing device of example 7, wherein the program instructions executable by the one or more processors such that the computing device is configured to obtain the third recording of the third audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration comprise program instructions executable by the one or more processors such that the computing device is configured to cause the computing device to obtain the third recording from the at least one microphone of the first playback entity.

[0364] Example 9: The computing device of example 7, wherein the program instructions executable by the one or more processors such that the computing device isconfigured to obtain the third recording of the third audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration comprise program instructions executable by the one or more processors such that the computing device is configured to cause the computing device to obtain the third recording from a microphone separate from the computing device.

[0365] Example 10: The computing device of example 9, wherein third recording of the third audio signal comprises one or more signals obtained from one or more microphones at different locations within the listening environment, wherein the one or more microphones at different locations within the listening environment comprise at least one or more of (i) a microphone at the computing device, (ii) a microphone at the first playback entity, (iii) a microphone at one of the one or more other playback entities in the area zone configuration, or (iv) a microphone separate from the computing device or the playback entities in the area zone configuration.

[0366] Example 11 : The computing device of example 1, wherein the computing device comprises one of (i) a computing device separate from the first playback entity and the one or more other playback entities in the area zone configuration, (ii) the first playback entity, or (iii) one of the one or more other playback entities in the area zone configuration.

[0367] Example 12: The computing device of example 1, wherein the program instructions further comprise program instructions executable by the one or more processors such that the computing device is configured to, for each additional playback entity of the one or more other playback entities in the area zone configuration: (i) cause the additional playback entity to emit a self-response audio signal via one or more speakers associated with the additional playback entity; (ii) obtain a recording comprising reflections of the self-response audio signal from one or more objects in the listening environment; and (iii) determine one or more boundary conditions of the additional playback entity based on the recording comprising the reflections of the self-response audio signal.

[0368] Example 13: The computing device of example 12, wherein the one or more additional playback entities in the area zone configuration comprise fewer than all of the playback entities in the area zone configuration, and wherein the program instructions comprise program instructions executable by the one or more processors such that the computing device is configured to, for the area zone configuration comprising the first playback entity and one ormore other playback entities located in the listening environment: cause a graphical user interface associated with the computing device to display a representation of the playback entities of the area zone configuration, and wherein the additional one or more playback entities correspond to playback entities selected via a user input received via the graphical user interface.

[0369] Example 14: A system comprising a first playback entity and one or more additional playback entities configured to play audio in a listening environment, wherein the first playback entity is configured to: (i) emit a first audio signal via one or more speakers associated with the first playback entity; (ii) obtain a first recording from at least one microphone associated with the first playback entity, wherein the first recording comprises reflections of the first audio signal from one or more objects in the listening environment; (iii) determine one or more boundary conditions of the first playback entity based on the first recording comprising the reflections of the first audio signal; (iv) emit a second audio signal via the one or more speakers associated with the first playback entity while each of the one or more additional playback entities in the system also emit the second audio signal, wherein the second audio signal is different than the first audio signal; (v) obtain a second recording of the second audio signal emitted by the first playback entity and each of the one or more additional playback entities in the system; (vi) determine a bass level in the listening environment based on the second recording; and (v) when the determined bass level in the listening environment differs from a target bass level for the listening environment by more than a threshold amount, adjust a bass equalization level of at least the first playback entity.

[0370] Example 15: The system of example 14, wherein the first audio signal comprises an ultrasonic audio signal, and wherein first playback entity is configured to: (i) determine whether one or more aspects of the reflections differ from one or more expected aspects of the reflections by more than a threshold amount; and (ii) when one or more aspects of the reflections differ from one or more expected aspects of the reflections by more than a threshold amount, adjust one or more equalization parameters of at least the first playback entity.

[0371] Example 16: The system of example 15, wherein the first playback entity is configured to: (i) determine whether one or more aspects of the reflections differ from one or more expected aspects of the reflections by more than a threshold amount by determining that the first playback entity’s proximity to at least one surface in the listening environment caused the one or more aspects of the reflections to differ from the one or more expected aspects of therefl ections by more than the threshold amount, wherein the at least one surface comprises one of a ceiling, a wall, or floor; and (ii) adjust one or more equalization parameters of the first playback entity based on the first playback entity’s proximity to the at least one surface in the listening environment.

[0372] Example 17: The system of example 14, wherein the first playback entity is configured to obtain the second recording of the second audio signal emitted by the first playback entity and each of the one or more additional playback entities in the system via at least one microphone integrated with the first playback entity.

[0373] Example 18: The system of example 17, wherein the first playback entity is configured to obtain the second recording of the second audio signal emitted by the first playback entity and each of the one or more additional playback entities in the system via a microphone separate from the first playback entity.

[0374] Example 19: The system of example 14, wherein the second audio signal comprises a broadband signal, and wherein the first playback entity is configured to determine the bass level in the listening environment based on the second recording based on an RT60 analysis of the second recording.

[0375] Example 20: The system of example 14, wherein the first playback entity is configured to: (i) emit a third audio signal via the one or more speakers while each of the one or more additional playback entities in the system also emit the third audio signal, wherein the third audio signal is different than the first audio signal and the second audio signal; (ii) obtain a third recording of the third audio signal emitted by the first playback entity and each of the one or more additional playback entities in the listening environment; (iii) determine an acoustic response of the listening environment based on the third recording; and (iv) when the determined acoustic response of the listening environment differs from a target acoustic response for the listening environment by more than a threshold amount, adjust one or more equalization levels of at least the first playback entity.

[0376] Example 21 : The system of example 20, wherein the first playback entity is configured to obtain the third recording of the third audio signal emitted by the first playback entity and each of the one or more additional playback entities in the system via at least one microphone integrated with the first playback entity.

[0377] Example 22: The system of example 20, wherein the first playback entity is configured to: obtain the third recording of the third audio signal emitted by the first playback entity and each of the one or more additional playback entities in the system via a microphone.

[0378] Example 23 :. The system of example 22, wherein third recording of the third audio signal comprises one or more signals obtained from one or more microphones at different locations within the listening environment, wherein the one or more microphones at different locations within the listening environment comprise at least one or more of (i) a microphone at a computing device separate from the first playback entity and the one or more additional playback entities in the system, (ii) a microphone at the first playback entity, (iii) a microphone at one of the one or more additional playback entities in the system, or (iv) a microphone separate from the computing device or the playback entities in the system.

[0379] Example 24: The system of example 14, wherein the one or more additional playback entities in the system comprises a second playback entity, wherein the second playback entity is configured to: (A) when each of the other playback entities in the system is not emitting an audio signal: (i) emit a self-response audio signal via one or more speakers associated with the second playback entity, (ii) obtain a reflection recording from at least one microphone associated with the second playback entity, wherein the reflection recording comprises reflections of the self-response audio signal from one or more objects in the listening environment, and (iii) determine a self-response of the second playback entity based on the reflection recording from at least one microphone associated with the second playback entity; (B) determine whether one or more aspects of the reflections in the reflection recording for the second playback entity differ from one or more expected aspects of the reflections for the second playback entity by more than a threshold amount; and (C) when one or more aspects of the reflections in the reflection recording for the second playback entity differ from one or more expected aspects of the reflections by more than a threshold amount, adjust one or more equalization parameters of at least the second playback entity.

[0380] Example 25: The system of example 24, further comprising tangible, non- transitory computer-readable media comprising program instructions executable by one or more processors to cause a computing device to perform functions comprising: (i) cause a graphical user interface associated with the computing device to display a representation of the playback entities of the system; and (ii) transmit data to and receive data from any one or of the playbackentities of the system based at least in part on one or more user inputs received via the graphical user interface.VIII. Conclusions

[0381] The above discussions relating to playback devices, controller devices, playback zone configurations, and media / audio content sources provide only some examples of operating environments within which functions and methods described below may be implemented. Other operating environments and configurations of media playback systems, playback devices, and network devices not explicitly described herein may also be applicable and suitable for implementation of the functions and methods.

[0382] The description above discloses, among other things, various example systems, methods, apparatus, and articles of manufacture including, among other components, firmware and / or software executed on hardware. It is understood that such examples are merely illustrative and should not be considered as limiting. For example, it is contemplated that any or all of the firmware, hardware, and / or software aspects or components can be embodied exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Accordingly, the examples provided are not the only ways) to implement such systems, methods, apparatus, and / or articles of manufacture.

[0383] Additionally, references herein to “embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one example embodiment of an invention. The appearances of this phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative example configurations mutually exclusive of other example configurations. As such, the example configurations described herein, explicitly and implicitly understood by one skilled in the art, can be combined with other example configurations.

[0384] The specification is presented largely in terms of illustrative environments, systems, procedures, steps, logic blocks, processing, and other symbolic representations that directly or indirectly resemble the operations of data processing devices coupled to networks. These process descriptions and representations are typically used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art. Numerous specific details are set forth to provide a thorough understanding of the present disclosure.However, it is understood to those skilled in the art that certain example configurations of the present disclosure can be practiced without certain, specific details. In other instances, well known methods, procedures, components, and circuitry have not been described in detail to avoid unnecessarily obscuring aspects of the example configurations. Accordingly, the scope of the present disclosure is defined by the appended claims rather than the foregoing description of example configurations.

[0385] When any of the appended claims are read to cover a purely software and / or firmware implementation, at least one of the elements in at least one example is hereby expressly defined to include a tangible, non-transitory medium such as a memory, DVD, CD, Blu-ray, and so on, storing the software and / or firmware.

Claims

CLAIMSWhat is claimed is:

1. A method for a computing device in communication with a first playback entity configured in a listening environment with at least one or more other playback entities, the method comprising: causing the first playback entity to emit a first audio signal via one or more speakers of the first playback entity; obtaining, via at least one microphone of the first playback entity, a first recording comprising reflections of the first audio signal from one or more objects in the listening environment; based on one or more boundary conditions of the first playback entity indicated by the first recording, causing the first playback entity to adjust an output of the first playback entity; causing the first playback entity to emit a second audio signal via the one or more speakers while each of the one or more other playback entities in an area zone configuration within the listening environment also emit the second audio signal; and based on a bass level in the listening environment indicated by a second recording of the second audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration, causing at least the first playback entity to adjust a bass equalization level of at least one playback entity in the area zone configuration.

2. The method of claim 1, further comprising: causing the first playback entity to emit an ultrasonic signal; and based on one or more aspects of detected reflections of the ultrasonic signal, causing the first playback entity adjusting one or more equalization parameters of the first playback entity.

3. The method of claims 1 or 2, wherein, when the reflections of the ultrasonic signal indicate that the one or more aspects of the detected reflections of the ultrasonic signal are caused by the first playback entity’s proximity to at least one surface in the listening environment, causing the first playback entity to adjust one or more equalization parameters ofthe first playback entity to at least partially compensate for the one or more aspects of the detected reflections of the ultrasonic signal.

4. The method of claim 3, wherein the at least one surface comprises one of a ceiling, a wall, or floor.

5. The method of any preceding claim, wherein the second recording is obtained from at least one of: at least one microphone of the first playback entity; and / or a microphone separate from the computing device.

6. The method of any preceding claim, wherein the second audio signal comprises a broadband signal, the method further comprising determining the bass level in the listening environment based on an RT60 analysis of the second recording.

7. The method of any preceding claim, further comprising: causing the first playback entity to emit a third audio signal via the one or more speakers while each of the one or more other playback entities in the area zone configuration also emit the third audio signal; obtain a third recording of the third audio signal emitted by the first playback entity and each of the one or more other playback entities in the area zone configuration; and based on an acoustic response of the listening environment indicated by the third recording, adjusting one or more equalization levels of at least one playback entity in the area zone configuration.

8. The method of claim 7, further comprising obtaining the third recording from at least one of: the at least one microphone of the first playback entity; and / or a microphone separate from the computing device.

9. The method of one of claims 7 and 8, wherein third recording of the third audio signal comprises one or more signals obtained from one or more microphones at different locations within the listening environment, wherein the one or more microphones at different locations within the listening environment comprise at least one or more of (i) a microphone at the computing device, (ii) a microphone at the first playback entity, (iii) a microphone at one of the one or more other playback entities in the area zone configuration, or (iv) a microphone separate from the computing device or the playback entities in the area zone configuration.

10. The method of any preceding claim 1, wherein the computing device comprises one of (i) a computing device separate from the first playback entity and the one or more other playback entities in the area zone configuration, (ii) the first playback entity, or (iii) one of the one or more other playback entities in the area zone configuration.

11. The method of any preceding claim, further comprising, for each additional playback entity of the one or more other playback entities in the area zone configuration: causing the additional playback entity to emit a self-response audio signal via one or more speakers associated with the additional playback entity; obtaining a recording comprising reflections of the self-response audio signal from one or more objects in the listening environment; and based on one or more boundary conditions of the additional playback entity indicated by the recording comprising the reflections of the self-response audio signal, causing the respective playback entity to adjust its output.

12. The method of any preceding claim, wherein the one or more other playback entities in the area zone configuration comprise fewer than all of the playback entities in the area zone configuration.

13. The method of claim 11, further comprising: causing a graphical user interface associated with the computing device to display a representation of the playback entities of the area zone configuration; andreceiving, via user input, an indication of an additional playback entity before causing the additional playback entity to emit the self-response signal.

14. A computing device comprising one or more processors configured to perform the method of any preceding claim.

15. A method for a first playback entity of a system comprising a first playback entity and one or more additional playback entities configured to play audio in a listening environment, the method comprising: emitting a first audio signal via one or more speakers associated with the first playback entity; adjusting an output of the first playback entity when a first recording from at least one microphone associated with the first playback entity indicates a boundary condition for the first playback entity, wherein the first recording comprises reflections of the first audio signal from one or more objects in the listening environment; emitting a second audio signal via the one or more speakers associated with the first playback entity while each of the one or more additional playback entities in the system also emit one or more audio signals; and adjusting a bass equalization level of at least the first playback entity based on a bass level in the listening environment indicated by a second recording of the second audio signal emitted by the first playback entity and the one or more audio signals emitted by each of the one or more additional playback entities in the system.

16. The method of claim 15, wherein the one or more audio signals emitted by the one or more playback entities in the system is the second audio signal.

17. The method of one of claims 15 to 16, wherein the first audio signal comprises an ultrasonic signal.

18. The method of one of claims 15 to 17, further comprising:determining whether one or more aspects of the reflections differ from one or more expected aspects of the reflections by more than a threshold amount; and when one or more aspects of the reflections differ from one or more expected aspects of the reflections by more than a threshold amount, adjust one or more equalization parameters of at least the first playback entity.

19. The method of one of claims 15 to 18, wherein the boundary condition corresponds to the first playback entity being proximate to at least one surface in the listening environment, thereby causing one or more aspects of the reflections to differ from the one or more expected aspects of the reflections.

20. The method of claim 19, wherein the at least one surface comprises one of a ceiling, a wall, or floor.

21. The method of one of claims 15 to 20, wherein the at least one microphone associated with the first playback entity is one of: integrated with the first playback entity; and separate from the first playback entity.

22. The method of one of claims 15 to 21, wherein the second audio signal comprises a broadband signal, wherein the bass level in the listening environment is determined based on the second recording based on an RT60 analysis of the second recording.

23. The method of one of claims 15 to 22, further comprising: emitting a third audio signal via the one or more speakers while each of the one or more additional playback entities in the system also emit the third audio signal, wherein the third audio signal is different than the first audio signal and the second audio signal; and adjust one or more equalization levels of at least the first playback entity based on an acoustic response of the listening environment determined based on a third recording of the third audio signal emitted by the first playback entity and each of the one or more additional playback entities in the listening environment.

24. The method of claim 23, wherein third recording of the third audio signal comprises one or more signals obtained from one or more microphones at different locations within the listening environment, wherein the one or more microphones at different locations within the listening environment comprise at least one or more of (i) a microphone at a computing device separate from the first playback entity and the one or more additional playback entities in the system, (ii) a microphone at the first playback entity, (iii) a microphone at one of the one or more additional playback entities in the system, or (iv) a microphone separate from the computing device or the playback entities in the system.

25. The method of one of claims 15 to 24, wherein the one or more additional playback entities in the system comprises a second playback entity, wherein the second playback entity is configured to: when each of the additional playback entities in the system is not emitting an audio signal: (i) emit a self-response audio signal via one or more speakers associated with the second playback entity, (ii) obtain a reflection recording from at least one microphone associated with the second playback entity, wherein the reflection recording comprises reflections of the selfresponse audio signal from one or more objects in the listening environment, and based on a selfresponse of the second playback entity determined based on the reflection recording from at least one microphone associated with the second playback entity, adjust one or more equalization parameters of at least the second playback entity.

26. A first playback entity configured to perform the method of one of claims 15 to 25.

27. A system comprising a first playback entity and one or more additional playback entities, the system configured to perform the method of one of claims 15 to 25.

Citation Information

Patent Citations

  • Immersive audio in a media playback system

    US10028069B1

  • Playback device calibration

    US10299061B1

  • Multiple-device setup

    US10303422B1

  • Media playback system with virtual line-in

    US10452345B1

  • Voice control of a media playback system

    US10499146B2