Devices and methods for learned device targeting

EP4714052A1Pending Publication Date: 2026-03-25SONOS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing media playback systems face challenges in accurately identifying target devices for interaction due to complex indoor environments and the inability to distinguish between devices based on signal strength alone, leading to potential user frustration and incorrect device selection.

Method used

The implementation of a parameterized machine learning model combined with BLUETOOTH Low Energy (BLE) signaling and logistic regression to predict user intent and target device selection, incorporating contextual influences and signal patterns to overcome environmental obstructions and ambiguities.

Benefits of technology

This approach enhances user experience by streamlining device interaction, reducing user effort, and improving accuracy in selecting the correct target device even in complex signaling environments, providing a more reliable and enjoyable media playback experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024029969_21112024_PF_FP_ABST
    Figure US2024029969_21112024_PF_FP_ABST
Patent Text Reader

Abstract

An example playback device is configured to detect beacon signals emitted by a control device, determine, for each beacon signal, an RSSI value and a standard deviation of a signal strength of the beacon signal to produce a first set of data, determine a first count of the beacon signals detected during a collection window, detect a plurality of reporting signals emitted by other playback devices, each reporting signal including a second set of data representing (i) a set of RSSI values and corresponding standard deviation values for a set of the beacon signals detected by the respective other playback device, and (ii) a second count of the set of the beacon signals during the collection window, and based on the first set of data, the first count, and the plurality of reporting signals, identify a proposed target playback device to receive audio playback instructions from the control device.
Need to check novelty before this filing date? Find Prior Art

Description

DEVICES AND METHODS FOR LEARNED DEVICE TARGETINGCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to co-pending U.S. Provisional Application No. 63 / 502,953 filed on May 18, 2023 and to co-pending U.S. Provisional No. 63 / 585,643 filed on September 27, 2023, each of which is hereby incorporated herein by reference in its entirety for all purposes.FIELD OF THE DISCLOSURE

[0002] The present disclosure is related to consumer goods and, more particularly, to methods, systems, products, aspects, services, and other elements directed to media playback or some aspect thereof.BACKGROUND

[0003] Options for accessing and listening to digital audio in an out-loud setting were limited until in 2002, when Sonos, Inc. began development of a new type of playback system. Sonos then filed one of its first patent applications in 2003, entitled “Method for Synchronizing Audio Playback between Multiple Networked Devices,” and began offering its first media playback systems for sale in 2005. The SONOS Wireless Home Sound System enables people to experience music from many sources via one or more networked playback devices. Through a software control application installed on a controller (e.g., smartphone, tablet, computer, voice input device), one can play what she wants in any room having a networked playback device. Media content (e.g., songs, podcasts, video sound) can be streamed to playback devices such that each room with a playback device can play back corresponding different media content. In addition, rooms can be grouped together for synchronous playback of the same media content, and / or the same media content can be heard in all rooms synchronously.BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Aspects, and advantages of the presently disclosed technology may be better understood with regard to the following description, appended claims, and accompanying drawings, as listed below. A person skilled in the relevant art will understand that the elements shown in the drawings are for purposes of illustrations, and variations, including different and / or additional elements and arrangements thereof, are possible.

[0005] FIG. 1 A is a partial cutaway view of an environment having a media playback system configured in accordance with aspects of the disclosed technology.

[0006] FIG. IB is a schematic diagram of the media playback system of FIG. 1A and one or more networks according to aspects of the disclosed technology.

[0007] FIG. 1C is a block diagram of a playback device according to aspects of the disclosed technology.

[0008] FIG. ID is a block diagram of a playback device according to aspects of the disclosed technology.

[0009] FIG. IE is a block diagram of a bonded playback device according to aspects of the disclosed technology.

[0010] FIG. IF is a block diagram of a network microphone device according to aspects of the disclosed technology.

[0011] FIG. 1G is a block diagram of a playback device according to aspects of the disclosed technology.

[0012] FIG. 1H is a partial schematic diagram of a control device according to aspects of the disclosed technology.

[0013] FIGS. II through IL are schematic diagrams of corresponding media playback system zones according to aspects of the disclosed technology.

[0014] FIG. IM is a schematic diagram of media playback system areas according to aspects of the disclosed technology.

[0015] FIG. 2 is a block diagram of one example of a positioning system that can be implemented in a media playback system, according to aspects of the disclosed technology.

[0016] FIG. 3 A is a plan view of another example of an environment having a media playback system configured according to aspects of the disclosed technology.

[0017] FIG. 3B is a plan view diagram of an example of the environment of FIG. 3 A showing signal transmissions according to aspects of the disclosed technology.

[0018] FIGS. 4 A through 4C are plan views of examples of the environment of FIG. 4 A showing signal transmissions according to aspects of the disclosed technology.

[0019] FIG. 5 is a flow diagram of one example of a process of target device prediction according to aspects of the disclosed technology.

[0020] FIG. 6 is a block diagram of one example of a machine learning personalization system according to aspects of the disclosed technology.

[0021] FIG. 7 is a flow diagram of one example of a learned device targeting process according to aspects of the disclosed technology.

[0022] FIG. 8 is a graph showing an example of a signal distribution according to aspects of the disclosed technology.

[0023] FIG. 9A is a plan view of one example of the environment of FIG. 3 A including a media playback system according to aspects of the disclosed technology.

[0024] FIG. 9B is a graph showing an example of a signal pattern corresponding to a location in the environment of FIG. 9 A.

[0025] FIG. 10 is a diagram of one example of a controller according to aspects of the disclosed technology.

[0026] FIG. 11A is a plan view of one example of the environment of FIG. 3 A including a media playback system according to aspects of the disclosed technology.

[0027] FIG. 1 IB is a graph showing examples of signal patterns corresponding to various locations in the environment of FIG. 11 A.

[0028] FIG. 12A is a plan view of one example of the environment of FIG. 3A including a media playback system according to aspects of the disclosed technology.

[0029] FIG. 12B is a graph showing examples of labeled signal patterns corresponding to locations in the environment of FIG. 9 A.

[0030] FIG. 12C is a plan view of another example of the environment and media playback system of FIG. 12A.

[0031] The drawings are for the purpose of illustrating example embodiments, but those of ordinary skill in the art will understand that the technology disclosed herein is not limited to the arrangements and / or instrumentality shown in the drawings.DETAILED DESCRIPTIONI. Overview

[0032] Embodiments described herein relate to techniques for personalizing a user experience with a media playback system and automatically identifying target devices for interaction based on predicted user intent. Many users demonstrate consistent listening routines or patterns when using capabilities and / or devices within their media playback system. By determining and recognizing consistent patterns over time, the media playback system can learn to predict certain user routines. In particular, as described further below, a parameterized machine learning model can be trained over time using positioning / localization information and passive user feedback to accurately predict a target device for interaction based on a location of the user’s control device and learned user preferences. By applying the techniques described herein, the system can reduce the time and user effort required to achieve the predicted endresult (e.g., the “time to music”) and provide more confidence in an easy and enjoyable experience as the number of interacting playback devices grows in a household.

[0033] Various user patterns and preferences may be strongly tied to location, and in some instances, to the proximity of the user to a particular playback device. As described further below, positioning / localization information can be obtained for a portable device (such as a portable playback device or a controller) based on patterns of wireless signals between the portable device and other devices in the media playback system. For example, signal strength measurements can provide an indication of proximity among devices. However, many (particularly indoor) environments are complex and contain numerous obstructions such that maximum signal strength alone may not provide an accurate indication of proximity. Further, in some instances, proximity without context may not correlate well with user intent. For example, a user operating their media playback system with a control device may be physically very close to a playback device that is in another room (separated from the user by a wall) but actually want to interact with a different playback device that is in the same room as the user, even though it is physically further away. Accordingly, aspects and embodiments provide techniques for incorporating contextual influence into the system’s predictions for device targeting, as described further below.

[0034] According to certain examples, BLUETOOTH Low Energy (BLE) signaling applied in combination with a parameterized machine learning model can be used to perform target device selection / identification. Approaches disclosed herein translate the concept of device targeting from purely proximity based sensing to user intent based targeting, giving users the opportunity to passively customize their experiences to fit their needs independently of the specific configuration (e.g., walls, furniture, etc.) of their environments. As discussed further below, certain examples provide techniques for applying logistic regression in the point-to-point signaling framework with BLE (or other) communication interfaces to achieve improved device targeting even in complex signaling environments (e.g., where wireless signals can travel through walls or other obstacles and / or multiple signal reflections may be present). These complex signaling environments can otherwise obscure correlations between user intent and the particular device(s) that are the target for a given action. Examples of the models and approaches disclosed herein have been shown to outperform methods that use only a maximum signal strength approach for target selection.

[0035] In some embodiments, for example, a playback device comprises a wireless communication interface configured to support communication of data via at least one network protocol, at least one processor, and at least one tangible non-transitory computer readablemedium storing program instructions that are executable by the at least one processor to cause the playback device to, during a collection window, detect, via the wireless communication interface, one or more beacon signals emitted by a control device, determine, for each beacon signal, a received signal strength indicator (RS SI) value and a standard deviation of a signal strength of the detected beacon signal relative to a median signal strength of the one or more beacon signals to produce a first set of measured data, and determine a first count of the one or more beacon signals detected during the collection window. The program instructions may further cause the playback device to detect, via the wireless communication interface, a plurality of reporting signals, each reporting signal emitted by a respective other playback device of a corresponding plurality of other playback devices, and each reporting signal including a second set of measured data representing (i) an RS SI value and corresponding standard deviation value for a set of the beacon signals detected by the respective other playback device during the collection window, and (ii) a second count of the set of the beacon signals detected by the respective other playback device during the collection window. Based on the first set of measured data, the first count, and the plurality of reporting signals, the playback device may identify from among a group including the playback device and the plurality of other playback devices, a proposed target playback device to receive, from the control device, playback instructions directing playback of audio content.

[0036] In further embodiments, for example, a media control device configured to control playback of audio content on a plurality of playback devices comprises a user interface, a wireless communication interface configured to support communication of data via at least one network protocol, and at least one processor. The media control device may further comprise at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the media control device to detect, via the user interface, a user input indicative of an intent to initiate playback of the audio content on at least one of the plurality of playback devices, to transmit, via the wireless communication interface, an instruction directing a first playback device of the plurality of playback devices to initiate a beaconing session among the plurality of playback devices, and to detect, via the wireless communication interface, one or more beacon signals emitted by one or more of the plurality of playback devices. For each of the one or more beacon signals, the media control device can be configured to determine a median RS SI value and a standard deviation of a signal strength relative to the median RS SI value to produce a set of measured data and to determine a count of the one or more beacon signals detected during the beaconing session. The media control device may further transmit to the first playback device, via the wireless communicationinterface, a reporting signal including the set of measured data and the count, detect, via the wireless communication interface, a message from the first playback device, the message identifying a proposed target playback device for playback of the audio content, and display, via the user interface, a suggestion to the user to select the proposed target playback device for playback of the audio content.

[0037] While some examples described herein may refer to functions performed by given actors such as “users,” “listeners,” and / or other entities, it should be understood that such references are for purposes of explanation only. The claims should not be interpreted to require action by any such example actor unless explicitly required by the language of the claims themselves.

[0038] In the Figures, identical reference numbers identify generally similar, and / or identical, elements. Many of the details, dimensions, angles, and other aspects shown in the Figures are merely illustrative of particular embodiments of the disclosed technology. Accordingly, other embodiments can have other details, dimensions, angles, and aspects without departing from the spirit or scope of the disclosure. In addition, those of ordinary skill in the art will appreciate that further embodiments of the various disclosed technologies can be practiced without several of the details described below.II. Suitable Operating Environment

[0039] FIG. 1A is a partial cutaway view of a media playback system (MPS) 100 distributed in an environment 101 (e.g., a house). In the illustrated embodiment of FIG. 1A, the environment 101 comprises a household having several rooms, spaces, and / or playback zones, including (clockwise from upper left) a master bathroom 101a, a master bedroom 101b, a second bedroom 101c, a family room or den 101 d, an office lOle, a living room 10 If, a dining room 101g, a kitchen lOlh, and an outdoor patio lOli. While certain embodiments and examples are described below in the context of a home environment, the technologies described herein may be implemented in other types of environments. In some embodiments, for example, the media playback system 100 can be implemented in one or more commercial settings (e.g., a restaurant, mall, airport, hotel, a retail or other store), one or more vehicles (e.g., a sports utility vehicle, bus, car, a ship, a boat, an airplane, etc.), multiple environments (e.g., a combination of home and vehicle environments), and / or another suitable environment where multi-zone audio may be desirable.

[0040] Within the rooms and spaces of the environment 101, the MPS 100 comprises one or more playback devices 110 (identified individually as playback devices HOa-n), one or morenetwork microphone devices 120 (“NMDs”) (identified individually as NMDs 120a-c), and one or more control devices 130 (identified individually as control devices 130a and 130b).

[0041] As used herein the term “playback device” can generally refer to a network device configured to receive, process, and output data of a media playback system. For example, a playback device can be a network device that receives and processes audio content. In some embodiments, a playback device includes one or more transducers or speakers powered by one or more amplifiers. In other embodiments, however, a playback device includes one of (or neither of) the speaker and the amplifier. For instance, a playback device can comprise one or more amplifiers configured to drive one or more speakers external to the playback device via a corresponding wire or cable.

[0042] Moreover, as used herein the term “NMD” (i.e., a “network microphone device”) can generally refer to a network device that is configured for audio detection. In some embodiments, an NMD is a stand-alone device configured primarily for audio detection. A stand-alone NMD 120 may omit components and / or functionality that is typically included in a playback device 110, such as a speaker or related electronics. For instance, in such cases, a stand-alone NMD may not produce audio output or may produce limited audio output. In other embodiments, an NMD is incorporated into a playback device (or vice versa). A playback device 110 that includes components and functionality of an NMD 120 may be referred to as being “NMD-equipped.” Examples of playback devices 110 and NMDs 120 are described further below.

[0043] The term “control device” can generally refer to a network device configured to perform functions relevant to facilitating user access, control, and / or configuration of the media playback system 100. Examples of control devices are described further below.

[0044] In some examples, one or more of the various playback devices 110 may be configured as portable playback devices, while others may be configured as stationary playback devices. For example, certain playback devices 110 may include an internal power source (e.g., a rechargeable battery) that allows the playback device to operate without being physically connected to a mains electrical outlet or the like. In this regard, such a playback device may be referred to herein as a “portable playback device.” On the other hand, playback devices that are configured to rely on power from a mains electrical outlet or the like may be referred to herein as “stationary playback devices,” although such devices may in fact be moved around a home or other environment. In practice, a person might often take a portable playback device to and from a home or other environment in which one or more stationary playback devices remain.

[0045] Each of the playback devices 110 is configured to receive audio signals or data from one or more media sources (e.g., one or more remote servers, one or more local devices, etc.) and play back the received audio signals or data as sound. The one or more NMDs 120 are configured to receive spoken word commands, and the one or more control devices 130 are configured to receive user input. In response to the received spoken word commands and / or user input, the media playback system 100 can play back audio via one or more of the playback devices 110. In certain embodiments, the playback devices 110 are configured to commence playback of media content in response to a trigger. For instance, one or more of the playback devices 110 can be configured to play back a morning playlist upon detection of an associated trigger condition (e.g., presence of a user in a kitchen, detection of a coffee machine operation, etc.). In some embodiments, for example, the media playback system 100 is configured to play back audio from a first playback device (e.g., the playback device 110a) in synchrony with a second playback device (e.g., the playback device 110b). Interactions between the playback devices 110, NMDs 120, and / or control devices 130 of the media playback system 100 configured in accordance with the various embodiments of the disclosure are described in greater detail below with respect to FIGS. 1B-1M.

[0046] The media playback system 100 can comprise one or more playback zones, some of which may correspond to the rooms in the environment 101. The media playback system 100 can be established with one or more playback zones, after which additional zones may be added, or removed, to form, for example, the configuration shown in FIG. 1 A. Each zone may be given a name according to a different room or space such as the office lOle, master bathroom 101a, master bedroom 101b, the second bedroom 101c, kitchen lOlh, dining room 101g, living room 10 If, and / or the balcony lOli. In some aspects, a single playback zone may include multiple rooms or spaces. In certain aspects, a single room or space may include multiple playback zones.

[0047] In the illustrated embodiment of FIG. 1A, the second bedroom 101c, the office lOle, the living room 10 If, the dining room 101g, the kitchen lOlh, and the outdoor patio lOli each include one playback device 110, and the master bathroom 101a, the master bedroom 101b, and the den 101 d include a plurality of playback devices 110. In the master bedroom 101b, the playback devices 1101 and 110m may be configured, for example, to play back audio content in synchrony as individual ones of playback devices 110, as a bonded playback zone, as a consolidated playback device, and / or any combination thereof. Similarly, in the den 101 d, the playback devices HOh-k can be configured, for instance, to play back audio content in synchrony as individual ones of playback devices 110, as one or more bonded playbackdevices, and / or as one or more consolidated playback devices. Additional details regarding bonded and consolidated playback devices are described below with respect to FIGS. IB, IE, and 1I-M.

[0048] In some aspects, one or more of the playback zones in the environment 101 may each be playing different audio content. For instance, a user may be grilling on the patio lOli and listening to hip hop music being played by the playback device 110c while another user is preparing food in the kitchen lOlh and listening to classical music played by the playback device 110b. In another example, a playback zone may play the same audio content in synchrony with another playback zone. For instance, the user may be in the office lOle listening to the playback device 1 lOf playing back the same hip hop music being played back by playback device 110c on the patio lOli. In some aspects, the playback devices 110c and 11 Of play back the hip hop music in synchrony such that the user perceives that the audio content is being played seamlessly (or at least substantially seamlessly) while moving between different playback zones. Additional details regarding audio playback synchronization among playback devices and / or zones can be found, for example, in U.S. Patent No. 8,234,395 titled, “System and method for synchronizing operations among a plurality of independently clocked digital data processing devices,” which is incorporated herein by reference in its entirety for all purposes. a. Suitable Media Playback System

[0049] FIG. IB is a schematic diagram of the media playback system 100 and a cloud network 102. For ease of illustration, certain devices of the media playback system 100 and the cloud network 102 are omitted from FIG. IB. One or more communication links 103 (referred to hereinafter as “the links 103”) communicatively couple the media playback system 100 and the cloud network 102.

[0050] The links 103 can comprise, for example, one or more wired networks, one or more wireless networks, one or more wide area networks (WAN), one or more local area networks (LAN), one or more personal area networks (PAN), one or more telecommunication networks (e.g., one or more Global System for Mobiles (GSM) networks, Code Division Multiple Access (CDMA) networks, Long-Term Evolution (LTE) networks, 5G communication networks, and / or other suitable data transmission protocol networks), etc. The cloud network 102 is configured to deliver media content (e.g., audio content, video content, photographs, social media content, etc.) to the media playback system 100 in response to a request transmitted from the media playback system 100 via the links 103. In some embodiments, the cloud network 102 is further configured to receive data (e.g., voice input data) from the media playback system100 and correspondingly transmit commands and / or media content to the media playback system 100.

[0051] The cloud network 102 comprises computing devices 106 (identified separately as a first computing device 106a, a second computing device 106b, and a third computing device 106c). The computing devices 106 can comprise individual computers or servers, such as, for example, a media streaming service server storing audio and / or other media content, a voice service server, a social media server, a media playback system control server, etc. In some embodiments, one or more of the computing devices 106 comprise modules of a single computer or server. In certain embodiments, one or more of the computing devices 106 comprise one or more modules, computers, and / or servers. Moreover, while the cloud network 102 is described above in the context of a single cloud network, in some embodiments the cloud network 102 comprises a plurality of cloud networks comprising communicatively coupled computing devices. Furthermore, while the cloud network 102 is shown in FIG. IB as having three of the computing devices 106, in some embodiments, the cloud network 102 comprises fewer (or more than) three computing devices 106.

[0052] The media playback system 100 is configured to receive media content from the networks 102 via the links 103. The received media content can comprise, for example, a Uniform Resource Identifier (URI) and / or a Uniform Resource Locator (URL). For instance, in some examples, the media playback system 100 can stream, download, or otherwise obtain data from a URI or a URL corresponding to the received media content. A network 104 communicatively couples the links 103 and at least a portion of the devices (e.g., one or more of the playback devices 110, NMDs 120, and / or control devices 130) of the media playback system 100. The network 104 can include, for example, a wireless network (e.g., a WI-FI network, a BLUETOOTH network, a Z-WAVE network, a ZIGBEE network, and / or other suitable wireless communication protocol network) and / or a wired network (e.g., a network comprising Ethernet, Universal Serial Bus (USB), and / or another suitable wired communication). As those of ordinary skill in the art will appreciate, as used herein, “WI-FI” can refer to several different communication protocols including, for example, Institute of Electrical and Electronics Engineers (IEEE) 802.11a, 802.11b, 802.11g, 802.1 In, 802.1 lac, 802.1 lac, 802.1 lad, 802.11af, 802.11 ah, 802.1 lai, 802.11aj, 802.1 laq, 802.1 lax, 802.1 lay, 802.15, etc. transmitted at 2.4 Gigahertz (GHz), 5 GHz, and / or another suitable frequency.

[0053] In some embodiments, the network 104 comprises a dedicated communication network that the media playback system 100 uses to transmit messages between individual devices and / or to transmit media content to and from media content sources (e.g., one or more of thecomputing devices 106). In certain embodiments, the network 104 is configured to be accessible only to devices in the media playback system 100, thereby reducing interference and competition with other household devices. In other embodiments, however, the network 104 comprises an existing household or commercial facility communication network (e.g., a household or commercial facility WI-FI network). In some embodiments, the links 103 and the network 104 comprise one or more of the same networks. In some aspects, for example, the links 103 and the network 104 comprise a telecommunication network (e.g., an LTE network, a 5G network, etc.). Moreover, in some embodiments, the media playback system 100 is implemented without the network 104, and devices comprising the media playback system 100 can communicate with each other, for example, via one or more direct connections, PANs, telecommunication networks, and / or other suitable communication links. The network 104 may be referred to herein as a “local communication network” to differentiate the network 104 from the cloud network 102 that couples the media playback system 100 to remote devices, such as cloud servers that host cloud services.

[0054] In some embodiments, audio content sources may be regularly added or removed from the media playback system 100. In some embodiments, for example, the media playback system 100 performs an indexing of media items when one or more media content sources are updated, added to, and / or removed from the media playback system 100. The media playback system 100 can scan identifiable media items in some or all folders and / or directories accessible to the playback devices 110, and generate or update a media content database comprising metadata (e.g., title, artist, album, track length, etc.) and other associated information (e.g., URIs, URLs, etc.) for each identifiable media item found. In some embodiments, for example, the media content database is stored on one or more of the playback devices 110, network microphone devices 120, and / or control devices 130.

[0055] In the illustrated embodiment of FIG. IB, the playback devices 1101 and 110m comprise a group 107a. The playback devices 1101 and 110m can be positioned in different rooms and be grouped together in the group 107a on a temporary or permanent basis based on user input received at the control device 130a and / or another control device 130 in the media playback system 100. When arranged in the group 107a, the playback devices 1101 and 110m can be configured to play back the same or similar audio content in synchrony from one or more audio content sources. In certain embodiments, for example, the group 107a comprises a bonded zone in which the playback devices 1101 and 110m comprise left audio and right audio channels, respectively, of multi-channel audio content, thereby producing or enhancing a stereo effect of the audio content. In some embodiments, the group 107a includes additional playback devices110. In other embodiments, however, the media playback system 100 omits the group 107a and / or other grouped arrangements of the playback devices 110. Additional details regarding groups and other arrangements of playback devices are described in further detail below with respect to FIGS. II through IM.

[0056] The media playback system 100 includes the NMDs 120a and 120b, each comprising one or more microphones configured to receive voice utterances from a user. In the illustrated embodiment of FIG. IB, the NMD 120a is a standalone device and the NMD 120b is integrated into the playback device HOn. The NMD 120a, for example, is configured to receive voice input 121 from a user 123. In some embodiments, the NMD 120a transmits data associated with the received voice input 121 to a voice assistant service (VAS) configured to (i) process the received voice input data and (ii) facilitate one or more operations on behalf of the media playback system 100.

[0057] In some aspects, for example, the computing device 106c comprises one or more modules and / or servers of a VAS (e.g., a VAS operated by one or more of SONOS, AMAZON, GOOGLE, APPLE, MICROSOFT, etc.). The computing device 106c can receive the voice input data from the NMD 120a via the network 104 and the links 103.

[0058] In response to receiving the voice input data, the computing device 106c processes the voice input data (i.e., “Play Hey Jude by The Beatles”), and determines that the processed voice input includes a command to play a song (e.g., “Hey Jude”). In some embodiments, after processing the voice input, the computing device 106c accordingly transmits commands to the media playback system 100 to play back “Hey Jude” by the Beatles from a suitable media service (e.g., via one or more of the computing devices 106) on one or more of the playback devices 110. In other embodiments, the computing device 106c may be configured to interface with media services on behalf of the media playback system 100. In such embodiments, after processing the voice input, instead of the computing device 106c transmitting commands to the media playback system 100 causing the media playback system 100 to retrieve the requested media from a suitable media service, the computing device 106c itself causes a suitable media service to provide the requested media to the media playback system 100 in accordance with the user’s voice utterance. b. Suitable Playback Devices

[0059] FIG. 1C is a block diagram of the playback device 110a comprising an input / output111. The input / output 111 can include an analog I / O I l la (e.g., one or more wires, cables, and / or other suitable communication links configured to carry analog signals) and / or a digital I / O 11 lb (e.g., one or more wires, cables, or other suitable communication links configured tocarry digital signals). In some embodiments, the analog I / O I l la is an audio line-in input connection comprising, for example, an auto-detecting 3.5mm audio line-in connection. In some embodiments, the digital I / O 111b comprises a Sony / Philips Digital Interface Format (S / PDIF) communication interface and / or cable and / or a Toshiba Link (TOSLINK) cable. In some embodiments, the digital I / O 111b comprises a High-Definition Multimedia Interface (HDMI) interface and / or cable. In some embodiments, the digital I / O 111b includes one or more wireless communication links comprising, for example, a radio frequency (RF), infrared, WI-FI, BLUETOOTH, or another suitable communication link. In certain embodiments, the analog I / O I l la and the digital I / O 111b comprise interfaces (e.g., ports, plugs, jacks, etc.) configured to receive connectors of cables transmitting analog and digital signals, respectively, without necessarily including cables.

[0060] The playback device 110a, for example, can receive media content (e.g., audio content comprising music and / or other sounds) from a local audio source 105 via the input / output 111 (e.g., a cable, a wire, a PAN, a BLUETOOTH connection, an ad hoc wired or wireless communication network, and / or another suitable communication link). The local audio source 105 can comprise, for example, a mobile device (e.g., a smartphone, a tablet, a laptop computer, etc.) or another suitable audio component (e.g., a television, a desktop computer, an amplifier, a phonograph (such as an LP turntable), a Blu-ray player, a memory storing digital media files, etc.). In some aspects, the local audio source 105 includes local music libraries on a smartphone, a computer, a networked-attached storage (NAS), and / or another suitable device configured to store media files. In certain embodiments, one or more of the playback devices 110, NMDs 120, and / or control devices 130 comprise the local audio source 105. In other embodiments, however, the media playback system omits the local audio source 105 altogether. In some embodiments, the playback device 110a does not include an input / output 111 and receives all audio content via the network 104.

[0061] The playback device 110a further comprises electronics 112, a user interface 113 (e.g., one or more buttons, knobs, dials, touch-sensitive surfaces, displays, touchscreens, etc.), and one or more transducers 114 (referred to hereinafter as “the transducers 114”). The electronics 112 are configured to receive audio from an audio source (e.g., the local audio source 105) via the input / output 111 or one or more of the computing devices 106a-c via the network 104 (FIG. IB), amplify the received audio, and output the amplified audio for playback via one or more of the transducers 114. In some embodiments, the playback device 110a optionally includes one or more microphones 115 (e.g., a single microphone, a plurality of microphones, a microphone array) (hereinafter referred to as “the microphones 115”). In certain embodiments,for example, the playback device 110a having one or more of the optional microphones 115 can operate as an NMD configured to receive voice input from a user and correspondingly perform one or more operations based on the received voice input.

[0062] In the illustrated embodiment of FIG. 1C, the electronics 112 comprise one or more processors 112a (referred to hereinafter as “the processors 112a”), memory 112b, software components 112c, a network interface 112d, one or more audio processing components 112g (referred to hereinafter as “the audio components H2g”), one or more audio amplifiers 112h (referred to hereinafter as “the amplifiers 112h”), and power 112i (e.g., one or more power supplies, power cables, power receptacles, batteries, induction coils, Power-over Ethernet (POE) interfaces, and / or other suitable sources of electric power). In some embodiments, the electronics 112 optionally include one or more other components 112j (e.g., one or more sensors, video displays, touchscreens, battery charging bases, etc.).

[0063] The processors 112a can comprise clock-driven computing component(s) configured to process data, and the memory 112b can comprise a computer-readable medium (e.g., a tangible, non-transitory computer-readable medium loaded with one or more of the software components 112c) configured to store instructions for performing various operations and / or functions. The processors 112a are configured to execute the instructions stored on the memory 112b to perform one or more of the operations. The operations can include, for example, causing the playback device 110a to retrieve audio data from an audio source (e.g., one or more of the computing devices 106a-c (FIG. IB)), and / or another one of the playback devices 110. In some embodiments, the operations further include causing the playback device 110a to send audio data to another one of the playback devices 110a and / or another device (e.g., one of the NMDs 120). Certain embodiments include operations causing the playback device 110a to pair with another of the one or more playback devices 110 to enable a multi-channel audio environment (e.g., a stereo pair, a bonded zone, etc.).

[0064] The processors 112a can be further configured to perform operations causing the playback device 110a to synchronize playback of audio content with another of the one or more playback devices 110. As those of ordinary skill in the art will appreciate, during synchronous playback of audio content on a plurality of playback devices, a listener will preferably be unable to perceive time-delay differences between playback of the audio content by the playback device 110a and the other one or more other playback devices 110. Additional details regarding audio playback synchronization among playback devices can be found, for example, in U.S. Patent No. 8,234,395, which is incorporated by reference above.

[0065] In some embodiments, the memory 112b is further configured to store data associated with the playback device 110a, such as one or more zones and / or zone groups of which the playback device 110a is a member, audio sources accessible to the playback device 110a, and / or a playback queue that the playback device 110a (and / or another of the one or more playback devices) can be associated with. The stored data can comprise one or more state variables that are periodically updated and used to describe a state of the playback device 110a. The memory 112b can also include data associated with a state of one or more of the other devices (e.g., the playback devices 110, NMDs 120, control devices 130) of the media playback system 100. In some aspects, for example, the state data is shared during predetermined intervals of time (e.g., every 5 seconds, every 10 seconds, every 60 seconds, etc.) among at least a portion of the devices of the media playback system 100, so that one or more of the devices have the most recent data associated with the media playback system 100.

[0066] The network interface 112d is configured to facilitate a transmission of data between the playback device 110a and one or more other devices on a data network such as, for example, the links 103 and / or the network 104 (FIG. IB). The network interface 112d is configured to transmit and receive data corresponding to media content (e.g., audio content, video content, text, photographs) and other signals (e.g., non-transitory signals) comprising digital packet data including an Internet Protocol (IP)-based source address and / or an IP -based destination address. The network interface 112d can parse the digital packet data such that the electronics 112 properly receive and process the data destined for the playback device 110a.

[0067] In the illustrated embodiment of FIG. 1C, the network interface 112d comprises one or more wireless interfaces 112e (referred to hereinafter as “the wireless interface 112e”). The wireless interface 112e (e.g., a suitable interface comprising one or more antennae) can be configured to wirelessly communicate with one or more other devices (e.g., one or more of the other playback devices 110, NMDs 120, and / or control devices 130) that are communicatively coupled to the network 104 (FIG. IB) in accordance with a suitable wireless communication protocol (e.g., WI-FI, BLUETOOTH, LTE, etc.). In some embodiments, the network interface 112d optionally includes a wired interface 112f (e.g., an interface or receptacle configured to receive a network cable such as an Ethernet, a USB-A, USB-C, and / or Thunderbolt cable) configured to communicate over a wired connection with other devices in accordance with a suitable wired communication protocol. In certain embodiments, the network interface 112d includes the wired interface 112f and excludes the wireless interface 112e. In some embodiments, the electronics 112 exclude the network interface 112d altogether and transmitand receive media content and / or other data via another communication path (e.g., the input / output 111).

[0068] The audio components 112g are configured to process and / or filter data comprising media content received by the electronics 112 (e.g., via the input / output 111 and / or the network interface 112d) to produce output audio signals. In some embodiments, the audio processing components 112g comprise, for example, one or more digital-to-analog converters (DACs), audio preprocessing components, audio enhancement components, digital signal processors (DSPs), and / or other suitable audio processing components, modules, circuits, etc. In certain embodiments, one or more of the audio processing components 112g can comprise one or more subcomponents of the processors 112a. In some embodiments, the electronics 112 omit the audio processing components 112g. In some aspects, for example, the processors 112a execute instructions stored on the memory 112b to perform audio processing operations to produce the output audio signals.

[0069] The amplifiers 112h are configured to receive and amplify the audio output signals produced by the audio processing components 112g and / or the processors 112a. The amplifiers 112h can comprise electronic devices and / or components configured to amplify audio signals to levels sufficient for driving one or more of the transducers 114. In some embodiments, for example, the amplifiers 112h include one or more switching or class-D power amplifiers. In other embodiments, however, the amplifiers 112h include one or more other types of power amplifiers (e.g., linear gain power amplifiers, class-A amplifiers, class-B amplifiers, class- AB amplifiers, class-C amplifiers, class-D amplifiers, class-E amplifiers, class-F amplifiers, class- G amplifiers, class H amplifiers, and / or another suitable type of power amplifier). In certain embodiments, the amplifiers 112h comprise a suitable combination of two or more of the foregoing types of power amplifiers. Moreover, in some embodiments, individual ones of the amplifiers 112h correspond to individual ones of the transducers 114. In other embodiments, however, the electronics 112 include a single one of the amplifiers 112h configured to output amplified audio signals to a plurality of the transducers 114. In some other embodiments, the electronics 112 omit the amplifiers 112h.

[0070] The transducers 114 (e.g., one or more speakers and / or speaker drivers) receive the amplified audio signals from the amplifier 112h and render or output the amplified audio signals as sound (e.g., audible sound waves having a frequency between about 20 Hertz (Hz) and 20 kilohertz (kHz)). In some embodiments, the transducers 114 can comprise a single transducer. In other embodiments, however, the transducers 114 comprise a plurality of audio transducers. In some embodiments, the transducers 114 comprise more than one type oftransducer. For example, the transducers 114 can include one or more low frequency transducers (e.g., subwoofers, woofers), mid-range frequency transducers (e.g., mid-range transducers, mid-woofers), and one or more high frequency transducers (e.g., one or more tweeters). As used herein, “low frequency” can generally refer to audible frequencies below about 500 Hz, “mid-range frequency” can generally refer to audible frequencies between about 500 Hz and about 2 kHz, and “high frequency” can generally refer to audible frequencies above 2 kHz. In certain embodiments, however, one or more of the transducers 114 comprise transducers that do not adhere to the foregoing frequency ranges. For example, one of the transducers 114 may comprise a mid-woofer transducer configured to output sound at frequencies between about 200 Hz and about 5 kHz.

[0071] By way of illustration, Sonos, Inc. presently offers (or has offered) for sale certain playback devices including, for example, a “SONOS ONE,” “PLAY:1,” “PLAY:3,” “PLAYA,” “PLAYBAR,” “PLAYBASE,” “CONNECT: AMP,” “CONNECT,” “AMP,” “PORT,” and “SUB.” Other suitable playback devices may additionally or alternatively be used to implement the playback devices of example embodiments disclosed herein. Additionally, one of ordinary skill in the art will appreciate that a playback device is not limited to the examples described herein or to Sonos product offerings. In some embodiments, for example, one or more playback devices 110 comprise wired or wireless headphones (e.g., over-the-ear headphones, on-ear headphones, in-ear earphones, etc.). In other embodiments, one or more of the playback devices 110 comprise a docking station and / or an interface configured to interact with a docking station for personal mobile media playback devices. In certain embodiments, a playback device may be integral to another device or component such as a television, an LP turntable, a lighting fixture, or some other device for indoor or outdoor use. In some embodiments, a playback device omits a user interface and / or one or more transducers. For example, FIG. ID is a block diagram of a playback device I lOp comprising the input / output 111 and electronics 112 without the user interface 113 or transducers 114.

[0072] FIG. IE is a block diagram of a bonded playback device 1 lOq comprising the playback device 110a (FIG. 1C) sonically bonded with the playback device HOi (e.g., a subwoofer) (FIG. 1A). In the illustrated embodiment, the playback devices 110a and 1 lOi are separate ones of the playback devices 110 housed in separate enclosures. In some embodiments, however, the bonded playback device HOq comprises a single enclosure housing both the playback devices 110a and HOi. The bonded playback device HOq can be configured to process and reproduce sound differently than an unbonded playback device (e.g., the playback device 110a of FIG. 1C) and / or paired or bonded playback devices (e.g., the playback devices 1101 and110m of FIG. IB). In some embodiments, for example, the playback device 110a is a full -range playback device configured to render low frequency, mid-range frequency, and high frequency audio content, and the playback device 1 lOi is a subwoofer configured to render low frequency audio content. In some aspects, the playback device 110a, when bonded with the first playback device, is configured to render only the mid-range and high frequency components of a particular audio content, while the playback device 1 lOi renders the low frequency component of the particular audio content. In some embodiments, the bonded playback device HOq includes additional playback devices and / or another bonded playback device. c. Suitable Network Microphone Devices (NMDs)

[0073] FIG. IF is a block diagram of the NMD 120a (FIGS. 1 A and IB). The NMD 120a includes one or more voice processing components 124 (hereinafter “the voice components 124”) and several components described with respect to the playback device 110a (FIG. 1C) including the processors 112a, the memory 112b, and the microphones 115. The NMD 120a optionally comprises other components also included in the playback device 110a (FIG. 1C), such as the user interface 113 and / or the transducers 114. In some embodiments, the NMD 120a is configured as a media playback device (e.g., one or more of the playback devices 110), and further includes, for example, one or more of the audio components 112g (FIG. 1C), the amplifiers 112h, and / or other playback device components. In certain embodiments, the NMD 120a comprises an Internet of Things (loT) device such as, for example, a thermostat, alarm panel, fire and / or smoke detector, etc. In some embodiments, the NMD 120a comprises the microphones 115, the voice processing components 124, and only a portion of the components of the electronics 112 described above with respect to FIG. 1C. In some aspects, for example, the NMD 120a includes the processor 112a and the memory 112b (FIG. 1C), while omitting one or more other components of the electronics 112. In some embodiments, the NMD 120a includes additional components (e.g., one or more sensors, cameras, thermometers, barometers, hygrometers, etc.).

[0074] In some embodiments, an NMD can be integrated into a playback device. FIG. 1G is a block diagram of a playback device 1 lOr comprising an NMD 120d. The playback device 1 lOr can comprise many or all of the components of the playback device 110a and further include the microphones 115 and voice processing components 124 (FIG. IF). The playback device HOr optionally includes an integrated control device 130c. The control device 130c can comprise, for example, a user interface (e.g., the user interface 113 of FIG. 1C) configured to receive user input (e.g., touch input, voice input, etc.) without a separate control device. Inother embodiments, however, the playback device 11 Or receives commands from another control device (e.g., the control device 130a of FIG. IB).

[0075] Referring again to FIG. IF, the microphones 115 are configured to acquire, capture, and / or receive sound from an environment (e.g., the environment 101 of FIG. 1A) and / or a room in which the NMD 120a is positioned. The received sound can include, for example, vocal utterances, audio played back by the NMD 120a and / or another playback device, background voices, ambient sounds, etc. The microphones 115 convert the received sound into electrical signals to produce microphone data. The voice processing components 124 receive and analyze the microphone data to determine whether a voice input is present in the microphone data. The voice input can comprise, for example, an activation word followed by an utterance including a user request. As those of ordinary skill in the art will appreciate, an activation word is a word or other audio cue signifying a user voice input. For instance, in querying the AMAZON VAS, a user might speak the activation word "Alexa." Other examples include "Ok, Google" for invoking the GOOGLE VAS and "Hey, Siri" for invoking the APPLE VAS.

[0076] After detecting the activation word, voice processing components 124 monitor the microphone data for an accompanying user request in the voice input. The user request may include, for example, a command to control a third-party device, such as a thermostat (e.g., NEST thermostat), an illumination device (e.g., a PHILIPS HUE lighting device), or a media playback device (e.g., a SONOS playback device). For example, a user might speak the activation word “Alexa” followed by the utterance “set the thermostat to 68 degrees” to set a temperature in a home (e.g., the environment 101 of FIG. 1 A). The user might speak the same activation word followed by the utterance “turn on the living room” to turn on illumination devices in a living room area of the home. The user may similarly speak an activation word followed by a request to play a particular song, an album, or a playlist of music on a playback device in the home. d. Suitable Control Devices

[0077] FIG. 1H is a partial schematic diagram of the control device 130a (FIGS. 1A and IB). As used herein, the term “control device” can be used interchangeably with “controller” or “control system.” Among other aspects, the control device 130a is configured to receive user input related to the media playback system 100 and, in response, cause one or more devices in the media playback system 100 to perform an action(s) or operation(s) corresponding to the user input. In the illustrated embodiment, the control device 130a comprises a smartphone (e.g., an iPhone™ an Android phone, etc.) on which media playback system controller applicationsoftware is installed. In some embodiments, the control device 130a comprises, for example, a tablet (e.g., an iPad™), a computer (e.g., a laptop computer, a desktop computer, etc.), and / or another suitable device (e.g., a television, an automobile audio head unit, an loT device, etc.). In certain embodiments, the control device 130a comprises a dedicated controller for the media playback system 100. In other embodiments, as described above with respect to FIG. 1G, the control device 130a is integrated into another device in the media playback system 100 (e.g., one more of the playback devices 110, NMDs 120, and / or other suitable devices configured to communicate over a network).

[0078] The control device 130a includes electronics 132, a user interface 133, one or more speakers 134, and one or more microphones 135. The electronics 132 comprise one or more processors 132a (referred to hereinafter as “the processors 132a”), a memory 132b, software components 132c, and a network interface 132d. The processor 132a can be configured to perform functions relevant to facilitating user access, control, and configuration of the media playback system 100. The memory 132b can comprise data storage that can be loaded with one or more of the software components executable by the processor 132a to perform those functions. The software components 132c can comprise applications and / or other executable software configured to facilitate control of the media playback system 100. The memory 132b can be configured to store, for example, the software components 132c, media playback system controller application software, and / or other data associated with the media playback system 100 and the user.

[0079] The network interface 132d is configured to facilitate network communications between the control device 130a and one or more other devices in the media playback system 100, and / or one or more remote devices. In some embodiments, the network interface 132d is configured to operate according to one or more suitable communication industry standards (e.g., infrared, radio, wired standards including IEEE 802.3, wireless standards including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G, LTE, etc.). The network interface 132d can be configured, for example, to transmit data to and / or receive data from the playback devices 110, the NMDs 120, other ones of the control devices 130, one of the computing devices 106 of FIG. IB, devices comprising one or more other media playback systems, etc. The transmitted and / or received data can include, for example, playback device control commands, state variables, playback zone and / or zone group configurations. For instance, based on user input received at the user interface 133, the network interface 132d can transmit a playback device control command (e.g., volume control, audio playback control, audio content selection, etc.) from the control device 130a to one or more of the playback devices110. The network interface 132d can also transmit and / or receive configuration changes such as, for example, adding / removing one or more playback devices 110 to / from a zone, adding / removing one or more zones to / from a zone group, forming a bonded or consolidated player, separating one or more playback devices from a bonded or consolidated player, among others. Additional description of zones and groups can be found below with respect to FIGS. II through IM.

[0080] The user interface 133 is configured to receive user input and can facilitate control of the media playback system 100. The user interface 133 includes media content art 133a (e.g., album art, lyrics, videos, etc.), a playback status indicator 133b (e.g., an elapsed and / or remaining time indicator), media content information region 133c, a playback control region 133d, and a zone indicator 133e. The media content information region 133c can include a display of relevant information (e.g., title, artist, album, genre, release year, etc.) about media content currently playing and / or media content in a queue or playlist. The playback control region 133d can include selectable (e.g., via touch input and / or via a cursor or another suitable selector) icons to cause one or more playback devices in a selected playback zone or zone group to perform playback actions such as, for example, play or pause, fast forward, rewind, skip to next, skip to previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit cross fade mode, etc. The playback control region 133d may also include selectable icons to modify equalization settings, playback volume, and / or other suitable playback actions. In the illustrated embodiment, the user interface 133 comprises a display presented on a touch screen interface of a smartphone (e.g., an iPhone™ an Android phone, etc.). In some embodiments, however, user interfaces of varying formats, styles, and interactive sequences may alternatively be implemented on one or more network devices to provide comparable control access to a media playback system.

[0081] The one or more speakers 134 (e.g., one or more transducers) can be configured to output sound to the user of the control device 130a. In some embodiments, the one or more speakers comprise individual transducers configured to correspondingly output low frequencies, mid-range frequencies, and / or high frequencies. In some aspects, for example, the control device 130a is configured as a playback device (e.g., one of the playback devices 110). Similarly, in some embodiments the control device 130a is configured as an NMD (e.g., one of the NMDs 120), receiving voice commands and other sounds via the one or more microphones 135.

[0082] The one or more microphones 135 can comprise, for example, one or more condenser microphones, electret condenser microphones, dynamic microphones, and / or other suitabletypes of microphones or transducers. In some embodiments, two or more of the microphones 135 are arranged to capture location information of an audio source (e.g., voice, audible sound, etc.) and / or configured to facilitate filtering of background noise. Moreover, in certain embodiments, the control device 130a is configured to operate as a playback device and an NMD. In other embodiments, however, the control device 130a omits the one or more speakers 134 and / or the one or more microphones 135. For instance, the control device 130a may comprise a device (e.g., a thermostat, an loT device, a network device, etc.) comprising a portion of the electronics 132 and the user interface 133 (e.g., a touch screen) without any speakers or microphones. e. Suitable Playback Device Configurations

[0083] FIGS. II through IM show example configurations of playback devices in zones and zone groups. Referring first to FIG. IM, in one example, a single playback device may belong to a zone. For example, the playback device 110g in the second bedroom 101c (FIG. 1A) may belong to Zone C. In some implementations described below, multiple playback devices may be “bonded” to form a “bonded pair” which together form a single zone. For example, the playback device 1101 (e.g., a left playback device) can be bonded to the playback device 110m (e.g., a right playback device) to form Zone B. Bonded playback devices may have different playback responsibilities (e.g., channel responsibilities). In another implementation described below, multiple playback devices may be merged to form a single zone. For example, the playback device 1 lOh (e.g., a front playback device) may be merged with the playback device 1 lOi (e.g., a subwoofer), and the playback devices 1 lOj and 110k (e.g., left and right surround speakers, respectively) to form a single Zone D. In another example, the playback devices 110b and 1 lOd can be merged to form a merged group or a zone group 108b. The merged playback devices 110b and HOd may not be specifically assigned different playback responsibilities. That is, the merged playback devices 110b and 1 lOd may, aside from playing audio content in synchrony, each play audio content as they would if they were not merged.

[0084] Each zone in the media playback system 100 may be provided for control as a single user interface (UI) entity. For example, Zone A may be provided as a single entity named Master Bathroom. Zone B may be provided as a single entity named Master Bedroom. Zone C may be provided as a single entity named Second Bedroom.

[0085] Playback devices that are bonded may have different playback responsibilities, such as responsibilities for certain audio channels. For example, as shown in FIG. II, the playback devices 1101 and 110m may be bonded so as to produce or enhance a stereo effect of audio content. In this example, the playback device 1101 may be configured to play a left channelaudio component, while the playback device 110m may be configured to play a right channel audio component. In some implementations, such stereo bonding may be referred to as “pairing.”

[0086] Additionally, bonded playback devices may have additional and / or different respective speaker drivers. As shown in FIG. 1 J, the playback device 1 lOh named Front may be bonded with the playback device 1 lOi named SUB. The Front device 1 lOh can be configured to render a range of mid to high frequencies and the SUB device 1 lOi can be configured to render low frequencies. When unbonded, however, the Front device I lOh can be configured to render a full range of frequencies. As another example, FIG. IK shows the Front and SUB devices 1 lOh and 1 lOi further bonded with Left and Right playback devices 1 lOj and 110k, respectively. In some implementations, the Left and Right devices HOj and 110k can be configured to form surround or “satellite” channels of a home theater system. The bonded playback devices 1 lOh, 1 lOi, 1 lOj, and 110k may form a single Zone D (FIG. IM).

[0087] Playback devices that are merged may not have assigned playback responsibilities and may each render the full range of audio content the respective playback device is capable of. Nevertheless, merged devices may be represented as a single UI entity (i.e., a zone, as discussed above). For instance, the playback devices 110a and HOn in the master bathroom have the single UI entity of Zone A. In one embodiment, the playback devices 110a and 1 lOn may each output the full range of audio content each respective playback devices 110a and 11 On are capable of, in synchrony.

[0088] In some embodiments, an NMD is bonded or merged with another device so as to form a zone. For example, the NMD 120b may be bonded with the playback device I lOe, which together form Zone F, named Living Room. In other embodiments, a stand-alone network microphone device may be in a zone by itself. In other embodiments, however, a stand-alone network microphone device may not be associated with a zone. Additional details regarding associating network microphone devices and playback devices as designated or default devices may be found, for example, in U.S. Patent No. 10,499,146 titled “Voice control of a media playback system” and filed on February 21, 2017, which is hereby incorporated herein by reference in its entirety for all purposes.

[0089] Zones of individual, bonded, and / or merged devices may be grouped to form a zone group. For example, referring to FIG. IM, Zone A may be grouped with Zone B to form a zone group 108a that includes the two zones. Similarly, Zone G may be grouped with Zone H to form the zone group 108b. As another example, Zone A may be grouped with one or more other Zones C-I. The Zones A-I may be grouped and ungrouped in numerous ways. Forexample, three, four, five, or more (e.g., all) of the Zones A-I may be grouped. When grouped, the zones of individual and / or bonded playback devices may play back audio in synchrony with one another, as described in previously referenced U.S. Patent No. 8,234,395. Playback devices may be dynamically grouped and ungrouped to form new or different groups that synchronously play back audio content.

[0090] In various implementations, the zones in an environment may be the default name of a zone within the group or a combination of the names of the zones within a zone group. For example, Zone Group 108b can be assigned a name such as “Dining + Kitchen”, as shown in FIG. IM. In some embodiments, a zone group may be given a unique name selected by a user.

[0091] Certain data may be stored in a memory of a playback device (e.g., the memory 112b of FIG. 1C) as one or more state variables that are periodically updated and used to describe the state of a playback zone, the playback device(s), and / or a zone group associated therewith. The memory may also include the data associated with the state of the other devices of the media system, and shared from time to time among the devices so that one or more of the devices have the most recent data associated with the system.

[0092] In some embodiments, the memory may store instances of various variable types associated with the states. Variable instances may be stored with identifiers (e.g., tags) corresponding to type. For example, certain identifiers may be a first type “al” to identify playback device(s) of a zone, a second type “bl” to identify playback device(s) that may be bonded in the zone, and a third type “cl” to identify a zone group to which the zone may belong. As a related example, identifiers associated with the second bedroom 101c may indicate that the playback device is the only playback device of the Zone C and not in a zone group. Identifiers associated with the Den may indicate that the Den is not grouped with other zones but includes bonded playback devices 11 Oh- 110k. Identifiers associated with the Dining Room may indicate that the Dining Room is part of the Dining + Kitchen zone group 108b and that devices 110b and 1 lOd are grouped (FIG. IL). Identifiers associated with the Kitchen may indicate the same or similar information by virtue of the Kitchen being part of the Dining + Kitchen zone group 108b. Other example zone variables and identifiers are described below.

[0093] In yet another example, the memory may store variables or identifiers representing other associations of zones and zone groups, such as identifiers associated with Areas, as shown in FIG. IM. An area may involve a cluster of zone groups and / or zones not within a zone group. For instance, FIG. IM shows an Upper Area 109a including Zones A-D and I, and a Lower Area 109b including Zones E-I. In one aspect, an Area may be used to invoke a cluster of zone groups and / or zones that share one or more zones and / or zone groups of another cluster. Inanother aspect, this differs from a zone group, which does not share a zone with another zone group. Further examples of techniques for implementing Areas may be found, for example, in U.S. Patent No. 10,712,997 filed August 21, 2017, and titled “Room Association Based on Name,” and U.S. Patent No. 8,483,853 filed September 11, 2007, and titled “Controlling and manipulating groupings in a multi-zone media system.” Each of these patents is incorporated herein by reference in its entirety. In some embodiments, the media playback system 100 may not implement Areas, in which case the system may not store variables associated with Areas. III. Positioning System Examples

[0094] As discussed above, a plurality of network devices, such as playback devices 110 and / or NMDs 120, can be distributed within an environment 101, such as a user’s home, or a commercial space (e.g., a restaurant, retail store, mall, hotel, etc.). Some of the devices may be in relatively fixed locations within the environment 101, whereas others may be portable and be frequently moved from one location to another. As the capabilities of these devices expand, it is becoming increasingly desirable to locate and interact with the devices within the environment 101. According to certain aspects, a positioning system can be implemented to determine relative positioning of devices within the environment 101 and optionally to control or modify behavior of one or more devices based on the relative positions. Positioning or localization information can be acquired through various techniques, optionally using sensors in some instances, examples of which are discussed below. In certain examples, one or more devices in the MPS 100, such as one or more playback devices 110, NMDs 120, or controller devices 130 may host a localization application that may implement operations (also referred to herein as functional capabilities or functionalities) that process localization information to enhance user experiences with the MPS 100. Examples of such operations include sophisticated acoustic manipulation (e.g., functional capabilities directed to psychoacoustic effects during audio playback) and autonomous device configuration and / or reconfiguration (e.g., functional capabilities directed to detection and configuration of new devices or devices that have moved or otherwise been changed in some way), among others. The requirements that these operations place on localization information vary, with some operations requiring low latency, high precision localization information and other operations being able to operate using high latency, low precision localization information.

[0095] According to certain examples, a positioning system can be implemented in the MPS 100 using a variety of different devices to generate the localization information utilized by certain application functionalities. However, the number, arrangement, and configuration of these devices can vary between examples. Additionally, or alternatively, the communicationstechnology and / or sensors employed by the devices can vary. Given the number of variables in play within any particular MPS and the concomitant inefficiencies that this variability imposes on MPS application operation development and maintenance, some examples disclosed herein utilize one or more playback devices 110, NMDs 120, or controller devices 130 to implement a positioning system using a common positioning application programming interface (API) that decouples the positioning / localization information from specific devices or underlying enabling technologies, as illustrated conceptually in FIG. 2.

[0096] Referring to FIG. 2, any one or more playback devices 110, NMDs 120, or controller devices 130 in the MPS 100 (“MPS devices”) can host a positioning system application 200. In certain implementations, one or more remote computing devices can facilitate hosting the application. The positioning system application 200 implements an application programming interface (API) that exposes positioning / localization information, and metadata pertinent thereto, to MPS application functionalities 202. The MPS functionalities 202 may include a wide variety of functional capabilities relating to various user experiences and aspects of the operation of the MPS 100. For example, the MPS functionalities 202 may include one or more VAS capabilities 204, such as voice disambiguation capabilities and arbitration between different NMDs receiving the same voice inputs, for example. The MPS functionalities 202 may also include one or more MPS and / or device configuration capabilities 206, such as automatic home theater configuration or reconfiguration, dynamically accommodating portable playback devices in home theater environments, dynamic room assignment for portable playback devices or their associated docks, and contextual orientation of controller devices 130, to name a few. The MPS functionalities 202 may further include one or more other functional capabilities 208 that use positioning / localization information. To support these and other MPS functionalities 202, positioning / localization information may be used to determine various pieces of information related to the locations of MPS devices within the environment 101. For example, the positioning / localization information may be used by some MPS functionalities 202 to keep track of which playback devices 110 or NMDs 120 are in a given room or space (e.g., which playback devices are in the Living Room 10 If, in which room is playback device HOd, or which playback devices 110 are closest to the controller device 130). The positioning / localization information may further be used to determine the distance and / or orientation between playback devices 110 (with varying levels of precision), or to determine the acoustic space around NMDs 120 or NMD-equipped playback devices 110 (e.g., which playback devices 110 can be heard from NMD 120a). Thus, the positioning / localization information may be used to determine information about the topology of the MPS 100 withinthe environment 101, which information may then be used to automatically and dynamically create or modify user experiences with the MPS 100 and support the MPS functionalities 202.

[0097] In some examples, the positioning / localization information is obtained through the exchange of wireless signals among network devices within the MPS 100. The positioning / localization information and metadata exposed by the positioning system application 200 may vary depending on the underlying communications technologies and / or sensor capabilities 210 within the MPS devices that are used to acquire the information and / or the needs of the particular MPS functionality 202. For example, certain MPS devices may be equipped with one or more network interfaces 224 that support any one or more of the following communications capabilities: Bluetooth 212, WI-FI 214 or ultra-wide-band technology (UWB 216; a short-range radio frequency communications technology). Further, certain MPS devices may be equipped to support signaling via acoustic signaling 218, ultrasound 220, or other signaling and / or communications means 222. Certain technologies 210 may be well-suited to certain MPS functionalities 202 while others may be more useful in other circumstances. For example, UWB 216 may provide high precision distance measurements, whereas WI-FI 214 (e.g., using RSSI signal strength measurements) or ultrasound 220 may provide “room-level” topology information (e.g., presence detection indicating that a particular MPS device is within a particular room or space of the environment 101). In some examples, combinations of the different technologies 210 may be used to enhance the accuracy and / or certainty of the information derived from the positioning / localization information received from one or more MPS devices via the positioning system application 200. For example, as discussed further below, in some instances, presence detection may be performed primarily using ultrasound 220; however, RSSI measurements may be used to confirm the presence detection and / or provide more precise localization information in addition to the presence detection.

[0098] Examples of MPS devices equipped with ultrasonic presence detection are disclosed in U.S. Patent Publication Nos. 2022 / 0066008 and 2022 / 0261212, each of which is hereby incorporated herein by reference in its entirety for all purposes. Examples of localizing MPS devices based on RSSI measurements are disclosed in U.S. Patent Publication No. 2021 / 0099736, which is herein incorporated by reference in its entirety for all purposes. Examples of performing location estimation of MPS devices using WI-FI 214 are disclosed in U.S. Patent Publication No. 2021 / 0297168, which is herein incorporated by reference in its entirety for all purposes.

[0099] In addition to the positioning / localization information itself, some examples of the positioning system application 200 can expose metadata that specifies localization capabilities of the host MPS device, such as precision and latency information and availability of the various underlying capabilities 210. As such, the positioning system application 200 enables the MPS functionalities 202 each to utilize a common set of API calls to identify the localization capability present within their host MPS device and to access positioning / localization information made available through the identified capabilities 210.

[0100] As shown in FIG. 2 and discussed above, the positioning system application 200 can interoperate with MPS devices that support a wide variety of localization capabilities, such as Bluetooth 212, WI-FI 214, UWB 216, acoustic signaling 218 and / or ultrasound 220, among others 222. In some examples, the positioning system application 200 includes one or more adapters configured to communicate with MPS devices using syntax and semantics specific to the localization capability 210 of the MPS devices. This architecture shields the MPS functionalities 202 from the complexity of interoperating with each type of MPS device. In some examples, each adapter can receive and process a stream of positioning / localization data from the MPS devices using any one or more of the communications capabilities 210. The adapters can interoperate with an accumulation engine within the positioning system application 200 that analyzes and merges (e.g., using a set of configurable rules) positioning / localization data obtained by the adapters and populates data structures that contain the positioning / localization information and the metadata described above. These data structures, in turn, are accessed and the positioning / localization information, and metadata, are retrieved by the positioning system application 200 in response to API calls received by the positioning system application 200 to support the MPS functionalities 202. The positioning / localization information, and metadata, can specify, in some examples, position / location of a device relative to other device, absolute position / location (e.g., within a coordinate system) of a device, presence of device (e.g., within a structure, room, or as a simple Boolean value), and / or orientation of a device.

[0101] For instance, in some examples, the positioning / localization information is expressed in two dimensions (e.g., as coordinates in a Cartesian plane), in three dimensions (e.g., as coordinates in a Cartesian space), or as coordinates within other coordinate systems. In certain examples, the positioning / localization information is stored in one or more data structures that include one or more records of fields typed and allocated to store portions of the information. For instance, in at least one example, the records are configured to store timestamps in association with values indicative of location coordinates of a portable playback device takenat a time given by the associated timestamp. Further, in at least one example, the records are configured to store timestamps in association with values indicative of a velocity of a portable playback device taken at a time given by the associated timestamp. Further, in at least one example, the records are configured to store timestamps in association with values indicative of a segment of movement (starting and ending coordinates) of a portable playback device taken at times given by associated timestamps. Other examples of positioning / localization information, and structures configured to store the same, will be apparent in view of this disclosure.

[0102] It should be noted that the API and adapters implemented by the positioning system application 200 may adhere to a variety of architectural styles and interoperability standards. For instance, in one example, the API is a web services interface implemented using a representational state transfer (REST) architectural style. In this example, the API communications are encoded in Hypertext Transfer Protocol (HTTP) along with JavaScript Object Notation and / or extensible markup language. In some examples, portions of the HTTP communications are encrypted to increase security. Alternatively, or additionally, in some examples, the API is implemented as a .NET web API that responds to HTTP posts to particular URLs (API endpoints) with localization data or metadata. Alternatively, or additionally, in some examples, the API is implemented using simple file transfer protocol commands. Also, in some examples, the adapters are implemented using a proprietary application protocol accessible via a user datagram protocol socket. Thus, the adapters and the API as described herein are not limited to any particular implementation.IV. Techniques for Learned Device Targeting

[0103] In many instances, user interactions with playback devices and / or certain device interactions with each other (e.g., forming a bonded group) are closely linked to the contextual location of devices within the environment 101. Although the positioning system described above can be used to determine the relative locations of MPS devices in the environment, numerous challenges remain with respect to location -based personalization. For example, due to wide variety in different floorplans and various interfering objects found indoors (e.g., walls, appliances, furniture, etc.), it can be difficult to reliably determine the location of a device in some environments. In addition, relative location or proximity alone may provide insufficient context to correctly identify a target device with which to interact. For example, a user may always turn on a playback device in the kitchen first thing in the morning from their bedroom, even though a closer playback device is present in their bedroom, because they intend to go to the kitchen. Similarly, if a portable playback device is moved into an area where there is ahome theater set-up, the user may wish to have the portable playback device form a bonded group with the home theater primary device, rather than with another device unrelated to the home theater set-up, even if that other device is physically closer to the portable playback device. Furthermore, while routines can play a significant role in users’ interactions with their media playback system, these routines can shift over time. For example, users may have a different routine during the week versus over the weekend, during the summer versus during the winter, or during school vacation periods versus during school semesters.

[0104] Accordingly, aspects and embodiments provide techniques for collecting household pattern data (e.g., device configuration settings, such as volume, playlist selection, etc., device movement within the environment, bonding information, etc.) to use in combination with positioning / localization information to train a target device prediction model specific to a user (or household). By incorporating this pattern data and applying logistic regression principles, examples can expand beyond simple proximity based targeting methods and allow position / location based device selection to be driven by usage history and user preferences. As described further below, certain examples employ implicit labeling to collected data, thereby allowing for continual learning and enhanced confidence in the model predictions. According to certain examples, labels for the data used in the device targeting model can be collected passively through user interaction by either correcting the target with a different selection or doing nothing and continuing with listening. In this manner, valuable user feedback can be acquired and used without requiring the user to actively perform calibration tasks or other actions outside of their normal interactions with the MPS 100.

[0105] In some instances, the arrangement of playback devices within an environment can change over time. For example, portable playback devices may be moved frequently from one location to another. In some examples, point-to-point signaling among playback devices and the collection of data samples for targeting allows for player movement and changing relative position to be identified by baseline reference pattern mismatch. As described further below, methods for anchor subset identification can be included for reuse of training data by the target device prediction model when encountering a new arrangement of devices. For example, certain stationary playback devices that may rarely change position within the environment can be identified as anchor devices that establish the baseline reference pattern for the environment. Training data associated with these anchor devices can be reused, even in cases where other, non-anchor devices (such as one or more portable devices) have moved. Furthermore, contextual location information, such as room detection, can be applied to link certain system behavior to specific locations. For example, if a portable playback device is moved into a roomcontaining a home theater set-up, and therefore may be a target device for bonding with one or more other home theater devices, the system can be configured to perform an automatic centering activity to prepare the portable playback device for becoming a home theater satellite. These and other examples are described further below.

[0106] Referring to FIG. 3A, there is illustrated an example of an environment 300 in which an MPS 100 can be deployed and in which various examples of the techniques and processes disclosed herein can be implemented. In this example, the environment 300 includes a plurality of rooms, namely, a master / primary bedroom 301a, a master bathroom 301b, an office 301c, a second bedroom 30 Id, a second bathroom 30 le, a living room 30 If, a dining room 301g, and a kitchen 301h. A plurality of playback devices 310 (individually identified as playback devices 310a-h) are distributed about the environment 300. The playback devices 310 may be any of the playback devices 110 or NMDs 120 described above. It will be appreciated that the layout of the environment 300 is intended as illustrative only, and a wide variety of other layouts and configurations are possible.

[0107] As described above, in certain examples, positioning / localization information can be obtained through the exchange of wireless signals among the playback devices 310. The wireless signals can be radio frequency (RF) signals, transmitted, for example, in accord with a BLE protocol or an 802.11 WI-FI protocol, or can be acoustic or ultrasound signals. Since the wireless signals attenuate with distance travelled, in some instances, measured signal strength at a receiving device can provide an indication of how close the transmitting device is to the receiving device. However, a multi-room environment, such as the environment 300, can present numerous challenges in terms of device localization and position-based control.

[0108] For any type of sensing, an obstruction leads to attenuation effects that have varying influences. Acoustic signals, for example, may be more strongly attenuated by obstructions than electromagnetic signals, such as radio transmissions. For example, 2.4 GHz radio transmissions (as may be used for BLE and some WI-FI signaling) are capable of penetrating the materials that make up typical homes, such as drywall, wood, glass, etc. As a result, walls and wood / glass furniture may not significantly attenuate these signals (depending on the thickness), whereas metal may have a stronger influence on signal strength. However, the nature, or even presence, of various obstructions within the environment may be unknown and as a result, signal strength often may not be a reliable indicator of proximity or distance between devices. Further, while the straight line distance between two playback devices 310 may have a strong influence on capabilities such as device targeting where RF signal is used, for acoustic signaling, obstructions may be more impactful than distance.

[0109] Different signaling technologies can have associated advantages and disadvantages. For example, since acoustic signals are sensitive to obstructions (including walls), this may limit the useful range of acoustic signaling for positioning, whereas acoustic signaling can be very useful for in-room presence detection. Further, certain playback devices 310 may be unavailable to transmit or receive acoustic tones, such as a device that is connected to a powered down TV or located within an acoustically sealed piece of furniture. The use of ultrasonic signaling may consume more power than using low energy BLUETOOTH (e.g., BLE) signaling, for example, which may be a significant consideration for battery constrained portable devices. Examples described below will refer to the use of BLE signals for positioning and device localization; however, it will be appreciated that the techniques described herein may be applied to signals other than BLE signals (e.g., WI-FI signals, acoustic signals, ultrasonic signals, etc.).

[0110] In certain examples, some or all of the playback devices 310 have a wireless communication interface (e.g., wireless interface 112e described above) that supports communication of data via at least one network protocol, such as a BLE and / or 802.11 WI-FI protocol, for example. Accordingly, the playback device 310 may include one or more WI-FI radios and / or BLE radios. A BLE radio may be configured to transmit and receive advertisement packets that allow for reading of the received signal strength indicator (RSSI) values associated with the BLE transmissions, which are used for location tracking and positioning, as described further below.[OHl] While the ability of certain radio signals (e.g., 2.4GHz transmissions) to penetrate walls and similar structures can have numerous advantages, it can also lead to ambiguity and confusion in the context of positioning and device targeting. For example, referring to FIG. 3B, in many areas within the environment 300, transmissions (represented by circles 302 surrounding various playback devices 310) from various playback devices 310 may overlap with one another. In many instances, these overlapping transmissions 302 can have the same or very similar signal strength. As a result, there can easily be misalignment between a desired target device for interaction and the target device that may be chosen based on maximum signal strength alone.

[0112] Referring again to FIG. 3 A, there are several scenarios where it can be difficult to differentiate a correct target device from among two or more potential target devices. For example, the ability of 2.4 GHz transmissions to penetrate walls can make it difficult to distinguish between playback devices in adjacent rooms. For example, from the point of view of a controller 330a in the second bedroom 30 Id, it can be ambiguous whether a target deviceis the playback device 310c in the second bedroom 301d or the nearby playback 310d, since the separating wall may not be a significant barrier to BLE positioning signals. In this case, the spatial proximity (through the wall) of the two playback devices 301c, 310d can cause ambiguity. Selection of the wrong target device due to this spatial positioning ambiguity may lead to a negative user experience. A similar issue is faced for environments with playback devices on multiple floors. In this case a playback device may be placed spatially close to (e.g., directly above or below) another, although on an adjacent floor. Accordingly, while a user may intend to target a playback device that is on the same floor as the user, signals from the playback device on the adjacent floor may be subject to the least total attenuation due to relative proximity and the attenuation imparted by the materials between the user’s control device and the two playback devices, such that the incorrect device is targeted.

[0113] In another example, contiguous open spaces, such as the living room 30 If, dining room 301g, and kitchen 301h, can also lead to ambiguity. For example, from the point of view of a controller 330b in the living room 301f, there is not a significant boundary to increase the amount of attenuation between potential target playback devices 310d, 310e, 31 Of, and 310g. As a result, there can be ambiguity as to which playback device 310 would be the correct target device in a given scenario. Thus, in these and other instances, received signal strength alone may not provide an accurate measure of contextual location or provide sufficient information for automated device targeting without significant potential for negative user experiences.

[0114] Aspects and embodiments provide techniques for addressing these issues and for providing a robust solution for automatic, personalized device targeting that can operate effectively in complex signaling environments, as may be encountered in numerous households and / or commercial settings. According to certain examples, there is provided a system that incorporates transmissions from a portable device, such as a controller 330 (e.g., a control device 130 described above), signaling between the playback devices 310, and user interactions that are used to train a parameterized machine learning model for device targeting. As described further below, the controller 330 and the playback devices 310 can be configured to transmit and detect BLE transmissions. These transmissions can be used to develop signal patterns that can be tied to various locations of the controller 330 within the environment 300 and / or associated with particular user behavior and preferences, including routine selection of one or more particular target playback devices 310. In this manner, repeated position / location-based activity can be learned and automated or streamlined for an enhanced user experience.

[0115] Certain approaches to determining playback device location, such as are described in U.S. Patent Publication No. 2021 / 0099736 referenced above, for example, involve the use ofpoint-to-point WI-FI signal transmissions between MPS devices and evaluate received signal strength indicator (RSSI) values. In some examples, received signals are normalized to account for device variations (e.g., one playback device 110 may employ a higher gain antenna than another) as well as environmental variations (e.g., reflected or multi-path signals, transmission through walls, etc.). These renormalized signal strengths can then be used to form a probabilistic framework to perform device targeting (e.g., determining which MPS device is the predicted target for interaction).

[0116] Aspects and examples disclosed herein extend these techniques to apply BLE signaling and logistic regression models to achieve robust, user intent driven target device prediction even in complex signaling environments, as described above. In particular, techniques disclosed herein provide ways to make user interaction with their multi-room audio system smoother, easier, more reliable, and delightful through its simplified use. As described above, one complication that may detract from user enjoyment is the confusion that can occur when navigating a controller application, such as the SONOS Controller App, on a control device (e.g., a smartphone or tablet) to identify the correct playback device to control. Similarly, complications can be encountered in the case of portable playback devices, such as a need for contextual location information to automate or streamline certain device activity (e.g., joining a bonded group or performing a centering process as described above). In all such cases, an improved understanding of a user’s intent through known location patterns may reduce friction associated with user interactions and lead to a more enjoyable experience.

[0117] According to certain examples, two signaling approaches can be employed to collect information that can be used to produce signal patterns that are in turn used by a machine learning personalization service to predict location-based actions, such as the selection of a target device for interaction. According to a first approach, signal collection between a portable device, such as the controller 330 or a portable playback device 310, for example, is used to establish a signal pattern that is tied to the location of the portable device. Referring to FIG. 4A, in some examples, the portable device employs passive signal collection, collecting “beacon” signals 304 that are broadcast by some or all of the other MPS devices in the network or environment 300. For the purposes of illustration, the following discussion will refer to the portable device as being the controller 330 (as illustrated in FIG. 4 A); however, it will be appreciated that the portable device may a device other than the controller 330, such as a portable playback device or another portable network device. FIG. 4A illustrates an example of the controller 330 in location 1 collecting beacon signals 304 emitted by playback devices310b, 310a, 310c, 3 lOd, 3 lOe, 31 Of, and 310g. A signal pattern can then be produced based on the collected beacon signals.

[0118] In other examples, the portable device (e.g. the controller 330) is further configured to emit its own beacon signals, as well as to detect beacon signals that are emitted by the other MPS devices. FIG. 4B illustrates an example of the controller 330 at location 2 capable of both transmitting and receiving beacon signals 304. In this case, some or all of the playback devices 310 produce signal data corresponding to those beacon signals 304 emitted by the controller 330 that the individual playback device 310 has detected. The signal data from the playback devices 310 can be collected and combined with signal data produced by the controller 330 (based on the beacon signals 304 it detected that were emitted by the other MPS devices), and the signal pattern can be produced based on the combined data.

[0119] In both of the above cases, the signal pattern is produced based on beacon signals 304 that are transmitted and / or received by the portable device. The resulting signal pattern is correlated with the location of the portable device at the time of the exchange of beacon signals, as described further below. For example, the signal pattern produced in the example of FIG. 4 will be correlated with location 1, whereas the signal pattern produced in the example of FIG. 4C will be correlated with location 2.

[0120] According to a second approach, the playback devices 310 can transmit and receive reference signals among themselves (point to point transmissions between playback devices). These reference signals can be used to produce a reference signal pattern that can be overlaid with the signal pattern produced based on the beacon signals and used to provide a reference framework for signal normalization and / or relative positioning of the portable device, as described further below. FIG. 4C illustrates such an example, with the controller 330 shown at location 3. In this example, the beacon signals 304 transmitted and / or received by the controller 330 and optionally one or more playback devices 310 are shown using dashed lines, while reference signals 306 transmitted and received among the playback devices 310 are shown using solid lines.

[0121] In certain examples, at least some of the playback devices 310 may be stationary devices that do not frequently change location in the environment 300. Accordingly, it will be appreciated that the reference pattern, or at least certain parts thereof, may remain relatively constant over time and be independent of the location of the controller 330. In contrast, each signal pattern produced based on the beacon signals 304 may be different depending on the corresponding location of the controller 330. Thus, the signal patterns produced from the beacon signals 304 can be used to by the machined learning personalization service to establishlocation-based personalization attributes, such as location-based device targeting, as described further below. Changes in the reference pattern can indicate changes in the environment 300, such as the addition of a new playback device 310 or movement of a playback device from one location to another. As such, these changes may be used as a trigger to update training of the model used by the personalization service since some learned location-based preferences may no longer be valid in the changed environment, as discussed further below.

[0122] Referring to FIG. 5, there is illustrated a flow diagram for one example of a process for learned device targeting in accord with various embodiments.

[0123] At operation 502, a portable device, such as a controller 330 or a portable playback device 310, can be configured to detect one or more beacon signals 304 emitted by one or more playback devices 310. In some examples, the beacon signals are BLE signals containing BLE advertisement packets. It will be appreciated that the beacon signals 304 may be collected from all or some playback devices 310, depending on various factors, including operational status of the individual playback devices 310 (e.g., whether or not a playback device is in a sleep mode), signaling capability of the individual playback devices 310 (e.g., whether or not a playback device has a BLE radio), locations of the controller 330 and the playback devices 310 within the environment 300, and / or the arrangement of the environment 300. In certain examples, the beacon signals 304 are collected during a beaconing session that may correspond to a predetermined time period, as described further below.

[0124] In some examples, the playback devices 310 may transmit multiple beacon signals 304 during the beaconing session. Thus, the controller 330 may detect far more beacon signals 304 than there are transmitting playback devices 310. Each of the collected beacon signals 304 may have a different signal strength, e.g., a different RSSI value, based on various factors, including the distance between the controller 330 and the source playback device 310 of the particular beacon signal 304 and the quantity and / or nature (e.g., material, thickness, etc.) of any obstacles in the path of the beacon signal 304. Accordingly, in some examples, the controller 330 may determine various characteristics of the beacon signals 304 as well as signal statistics during the beaconing session. For example, the controller 330 may determine the RSSI value for each beacon signal, the median signal strength (or RSSI value) for the group of collected beacon signals 304, the standard deviation of the signal strength for each collected beacon signal relative to the median signal strength, and a count of the total number of beacon signals 304 detected during the beaconing session.

[0125] In examples, each of the beacon signals 304 includes identification information that identifies the particular playback device 310 that is the source of the respective beacon signal304. For example, the beacon signals 304 may each include a sequence of tones and / or a transmission identifier. In some examples, the sequence of tones is specific for each playback device 310 and can therefore be used to identify the playback device that is the source of the beacon signal 304. In other examples, the transmission identifier identifies the playback device 310 that is the source of the beacon signal 304. In such examples, the controller 330 may group the detected beacon signals according to the source playback device from which they originated, and determine signal statistics for each group of beacon signals.

[0126] As discussed above, in some examples, signal collection during the beaconing session at operation 502 is performed only by the controller 330 (passive signal collection), as in the example of FIG. 4A. In other examples, signal collection during the beaconing session at operation 502 is performed by the controller 330 and one or more of the playback devices 310, as in the example of FIG. 4B. Accordingly, in such examples, one or more of the playback devices 310 can be configured to collect, during operation 502, the beacon signals emitted by the controller 330 and to produce signal measurement data based thereon. In some examples, the controller 330 emits multiple beacon signals during a given beaconing session. Accordingly, the playback device(s) 310 may determine RS SI values for each collected beacon signal, along with signal statistics, such as the median RS SI value for the group of collected beacon signals, the standard deviation of the signal strength for each collected beacon signal relative to the median signal strength, and a count of the total number of beacon signals detected during the beaconing session, for example. The signal measurement data collected by the one or more playback devices 310 can be combined with the signal measurement data collected by the controller 330 and the combined data sets can be used at operation 504 to develop the signal pattern. Examples of this process are described further below.

[0127] Based on the information collected and determined during the beaconing session, at operation 504, a signal pattern corresponding to that beaconing session, and therefore to the location of the controller 330 during the beaconing session, can be established. In some examples, the time duration of the beaconing session can be selected such that movement of the controller 330 during the beaconing session is likely to be minimal (e.g., a few milliseconds or up to one or two seconds). Various other factors may also influence the selection of the time duration of the beaconing session, as described further below. The signal pattern includes both signal information (e.g., RSSI values and statistics) and playback device information (e.g., which playback devices 310 contributed beacon signals 304 to the pattern).

[0128] The signal pattern developed at operation 504 may be unique to, or at least strongly tied to or dependent on, the corresponding location of the controller 330. For example, the signalpattern that may be developed for the controller 330 in location 1 as shown in FIG. 4 A may be identifiably different from a signal pattern that may be developed for the controller 330 in location 2 as shown in FIG. 4B or location 3 as shown in FIG. 4C.

[0129] As discussed above, in certain examples, at operation 506, a reference pattern is produced based on reference signals 306 exchanged among the playback devices 310. In some examples, the reference signals 306 are BLE signals, similar to or the same as the beacon signals 304. In other examples, a different signaling technology or protocol can be used for transmission and reception of the reference signals 306. For example, the reference signals 306 may be acoustic signals or ultrasound signals. In some examples, the reference signals 306 are transmitted and received, and the reference pattern is produced at operation 506 during the beaconing session or during a time overlapping with the beaconing session. In other examples, the reference signals 306 can be transmitted and received, and the reference pattern can be produced, independent of the beaconing session.

[0130] At operation 508, a particular activity, such as a playback session being activated on a particular playback device 310 or a particular playback device 310 being added to a bonded group, for example, is linked to the corresponding signal pattern developed at operation 504. Accordingly, in some examples, a beaconing session at operation 502 can be initiated by a user’s interaction with the controller 330, thus indicating to the system that an activity is about to occur. Thus, the resulting activity that triggered the start of a beaconing session can be reliably linked to the signal pattern that corresponds to that beaconing session. For example, referring to FIG. 4A, if a user, through interaction with the controller 330 (e.g., via the user interface 133) begins a playback session using a bonded group of playback devices including the playback devices 3 lOd, 3 lOe, and 31 Of, this activity initiates a beaconing session that allows the signal pattern corresponding to location 1 of the controller 330 in FIG. 4 A to be produced. Accordingly, this activity can be linked to the corresponding signal pattern and location 1. In particular, the playback devices 30 Id, 3 lOe, and 3 lOf can be identified as the target devices for this activity. Similarly, referring to FIG. 4B, if a user, through interaction with the controller 330, begins a playback session on the playback device 310h, this activity can be associated with the signal pattern corresponding to location 2 of the controller 330 in FIG. 4B. In particular, in this example, the playback device 3 lOh can be identified as the target device.

[0131] As described above, in some examples, the reference pattern produced at operation 506 can be used to localize the position of the controller 330 relative to one or more playback devices 310. Accordingly, in such examples, at operation 510, the signal pattern produced atoperation 504 can also be linked to the corresponding relative position / location of the controller 330.

[0132] The signal patterns and the linked activities can be used as inputs for a parameterized machine learning model. For example, the beacon signal measurements (e.g., RSSI values, signal statistics, etc.) can be used as model features, while the identities of the target devices can be used as labels for the corresponding signal pattern data, as described further below. Accordingly, at operation 512, these inputs can be used to train, and then once trained, apply, a device targeting machine learning model to predict target devices based on recognizing signal patterns produced during future beaconing sessions.

[0133] At operation 512, in response to a received signal pattern developed at operation 504, the trained model outputs the identity of a predicted target device. Examples and operation of a device targeting model are described further below with reference to FIGS. 6 and 7.

[0134] As described above, according to certain examples, the MPS 100 can be configured to implement a personalization service that incorporates various machine learning approaches to personalize one or more attributes of the MPS, including automated target device prediction. The personalization service can be implemented by one or more of the playback devices 310 and / or the controller 308, individually or in combination. In some examples, personalization functionality, including learned device targeting, can be accomplished using a model predictive controller that runs a parameterized machine learning model.

[0135] FIG. 6 illustrates an example of a machine learning system 600 that can be used to implement learned device targeting, optionally as well as other personalization functionality. In this example, a model predictive controller 602 operates on input data 604 and its operation is controlled by an optimizer 606. The model predictive controller (MPC) 602 includes a model 608, which may be a parameterized machine learning model, as discussed above. The MPC 602 further includes a data sampler 610, user preferences 612, and a confidence element 614. The confidence element 614 may apply two threshold values, namely a decision threshold 616 and an uncertainty threshold 618, each of which is discussed further below. The confidence element 614 allows the system to accommodate uncertainty in the prediction (e.g., by using confidence indicators, as described below), which can lead to improved performance. The system 600 may be implemented, in whole or in part, on one or more network devices (e.g., playback devices 310 or controller 330) within the MPS 100, or may be implemented, in whole or in part, on a cloud network device 102, for example. The system 600 may be implemented in software or using any combination of hardware and software capable of performing the functions disclosed herein.

[0136] In examples, the model predictive controller 602 runs the model 608 based on parameters associated with one or more features extracted from the input data 604 to produce a personalization result or recommendation, such as a predicted target device for a given interaction. The input data 604 can comprise any data which is used to correlate user behavior with a specific action and target device. As described above, in some examples, the input data 604 includes the signal measurement data that is collected by the controller 330 and / or one or more playback devices 310 during beaconing sessions. The input data 604 may further include the activity and target device identity associated with a signal pattern produced for each beaconing session. Accordingly, the input data 604 may be “local” data in that it is data collected from a specific environment 300. In some examples, particularly to assist the system 600 when little or no local data is available (e.g., when the system 600 is first activated or reactivated after a long period of inactivity), the input data 604 may include some “global” data that is data collected from outside sources, such as a group of one or more other environments 300 or averaged trends from multiple environments 300, for example.

[0137] In examples, the data sampler 610 intakes the input data 604 and extracts one or more input features to be used by the model 608, as described further below. In some examples in which the input data 604 includes both local data and global data, the data sampler 610 determines how to combine the local and global data. In some examples, this operation of the data sampler 610 can be modified by a “proportion” hyperparameter that determines the mix, or by a more sophisticated sampling regime, for example.

[0138] The model 608 uses the input data 604 to generate a set of parameters which yield a generalized function capable of predicting one or more particular output values (e.g., the identifier of a target device) based on new input data 604. Parameters are variables that “belong” to the model 608 in that the trained model is represented by the model parameters. In contrast, hyperparameters are higher-level variables that affect the learning process and, thus, the values of the model parameters of the trained model 608. In some examples, training the model 608 involves choosing hyperparameters that the learning process uses to generate parameters that correctly map the input features (independent variables) to the labels (dependent variable) such that the model 608 produces predictions (e.g., target device identities) with reasonable accuracy.

[0139] In the example illustrated in FIG. 6, the system 600 includes the optimizer 606 that operates based on one or more hyperparameters to optimize performance of the MPC 602. In some examples, the optimizer 606 selects hyperparameters for use during training of the model 608. Hyperparameters may include variables that determine characteristics such as anarchitecture of the model 608 (e.g., kernel selection or type of model (e.g., linear regression, Gaussian process, logistic regression, gradient boosted tree classifier, etc.), kernel size, etc.) how the model 608 is applied, the mix of local and / or global input data used, and / or variables that affect an optimization process used by the optimizer 606. A hyperparameter can, for example, take the form of a single continuous scalar variable or a discrete categorical variable (e.g., which kernel to use). Selection of hyperparameters has a significant impact on the performance (e.g., accuracy of predictions) of the trained model 608. Accordingly, in some examples, the optimizer 606 applies an optimization process to select the best hyperparameters for training the model 608. In some examples, this optimization process involves testing the performance of the system 600 on a validation dataset and adjusting the hyperparameters to produce an optimal result. In some examples, the optimizer 606 may use a grid search involving a field of combinations of hyperparameter values. In other examples, the optimizer 606 may apply a gradient descent optimization or a gradient-free optimization method, such as Bayesian optimization, or some combination thereof. As noted above, in some examples, the choice of optimization process can be a hyperparameter itself.

[0140] Thus, hyperparameters are “external” to the model 608 since they cannot be changed by the model during training, although they are tuned by the optimizer 606 to control the training of the model 608. As described above, a hyperparameter selected by the optimizer 606 can include a set of model parameters, as well as values that define the model architecture itself. In contrast, the model parameters are internal to the model 608 and their values are learned or estimated based on the input data 604 during training as the model 608 tries to learn the mapping between the input features and the labels. In some examples, training of the model 608 begins with the parameter values set to some initial default values (e.g., random values or set to zeros), and these initial values are updated as training / learning progresses under control of the optimizer 606.

[0141] According to certain examples, the model 608 is configured to find the probability, P, of a label, y, given a set of features, x, with parameters, 0, for each ithmeasurement according to the function:In the above function, Fl, o is a logic cost function that is defined by the model architecture. In some examples, the model 608 is selected to be a logistic regression model, and accordingly, c is given by:In this example, there are N features (x) each with a parameter (9) that is fit through minimization of the cost function.

[0142] Still referring to FIG. 6, as discussed above, in some examples, the MPC 602 incorporates uncertainty through the use of the confidence element 614. In the above-discussed example, the output from the model 608 is a probability and therefore has a built-in measure of uncertainty, or “confidence metric.” For example, in the case of learned device targeting, the output from the model 608 is a probability of a particular target device identity based on the input data 604 corresponding to the current beaconing session. In one example, the uncertainty threshold 618 dictates a value at which the uncertainty in the model output is sufficiently low for the MPC 602 to trust the model prediction. For example, if the model output (prediction) is a 60% probability that the target device is the playback device 310b, the uncertainty (40% in this case) may be too high for the MPC 602 to trust the model prediction. This may indicate a need for re-training of the model 608, for example. In some examples, the decision threshold 616 dictates a value at which the uncertainty in the model output is sufficiently low for the MPC 602 to take a certain action based on the model output. This action may include providing a suggestion of the target device to the user, or even automatically selecting the target device.

[0143] To minimize friction with a user and avoid negative user experiences, the decision threshold may be set at a relatively high value (very low uncertainty) and may be significantly higher than the uncertainty threshold. For example, while a probability of 90% that the target device is playback device 310b may be sufficiently low uncertainty to indicate that the model 608 is operating correctly, the uncertainty may still be too high for the MPC 602 to autonomously direct selection of playback device 310b as the target device. In this case, the MPC 602 may suggest the target device to the user or take no action with respect to the model prediction.

[0144] In some instances, the uncertainty threshold 608 and / or the decision threshold 616 can be hyperparameters that are applied (and optionally tuned) by the optimizer 606. For example, the uncertainty threshold 608 and / or the decision threshold 616 may together define a trust region in which it is likely that acting on the model prediction will not result in undesirable system behavior (e.g., selecting the wrong target device) and negative user experience. In some examples, the optimizer 606 can be constrained to optimize the model parameters within this trust region set by the uncertainty threshold 608 and the decision threshold 616. In other examples, the uncertainty threshold 608 and / or the decision threshold 616 may not be used to tune the model parameters during training, but may instead directly affect the decision behaviorof the MPC 602. For example, as described above, in some instances, the MPC 602 can be configured to automatically take an action (such as selecting a target device) if the uncertainty in the model prediction is below the limit set by the decision threshold (e.g., below 10%, 5%, or 2% uncertainty, etc.). In some examples, the MPC may offer a suggestion of the target device to the user if the uncertainty in the model prediction is below the limit set by the decision threshold 616, or is above the limit set by the decision threshold 616 but below the limit set by the uncertainty threshold 608 (e.g., within the trust region). Various other scenarios will be apparent given the benefit of this disclosure. Thus, the confidence element 614 can provide a valuable resource in terms of configuring the system 600 to provide useful target device recommendations to users and reduce instances of providing incorrect, unwanted, or annoying suggestions or actions.

[0145] According to certain examples, the MPC 602 may further acquire and store user preference information, as indicated at 612. The user preference information may include user- provided information regarding the level of personalization desired by the user, playback device attributes or configurations that the user does or does not want to be personalized (e.g., a user may agree to automatic target device selection in some scenarios, but not others, or may indicate that while target device suggestions may be provided, automatic target device selection is not permitted), and / or other user preferences with respect to personalization functionalities described herein. The user preference information may be acquired as part of the input data 604 in some examples or may be separately acquired and stored. In some examples, a user may enter the user preferences via a user interface 620, such as the user interface 133 on a control device 130, for example. In some examples, the user preferences 612 can be used to control how the hyperparameters are selected. For example, the optimizer 606 can be configured to optimize the model parameters within constraints set by the user preferences 612. User preferences 612 may also directly influence the behavior of the MPC 602, such as by constraining automated actions to certain time periods, and / or scenarios, or by forbidding automatic action (e.g., automatic target device selection) and allowing suggestions only. In some examples, the user preferences can be used to set either or both of the decision threshold 616 and / or the uncertainty threshold 618. In this manner, a user can be provided with a wide degree of control over the behavior of the system 600 such that the system 600 can be configured in accord with an individual user’s own preferences and comfort level with system autonomy and personalization.

[0146] Furthermore, passive user feedback can be used to gauge the accuracy of the model predictions and adjust the model 608 if necessary to improve performance. For example, if thesystem 600 suggests a target device to the user via the user interface 620 and the user selects that target device, the MPC 602 may interpret that the predicted target device was correct. On the other hand, if the user selects a different target device, the MPC 602 may interpret that the prediction was incorrect. As described further below, this passive user feedback can be used to label the corresponding features associated with the signal measurement data set that produced the prediction and produce labeled training data. By re-training the model 608 with this labeled training data, the model 608 may produce similar predictions with higher or lower confidence metrics. In some examples, where the labeled training data is based on positive user feedback, the re-trained model 608 may produce a corresponding target device prediction based on similar input data 604 with a higher confidence metric (e.g., higher probability). In other examples, where the labeled training data is based on negative user feedback, the model may be less likely to produce a corresponding target device prediction based on similar input data 604, or if it does, may produce the corresponding target device prediction with a lower confidence metric (e.g., lower probability).

[0147] FIG. 7 is a flow diagram illustrating one example of a target device prediction / selection process that may be implemented by the MPS 100 using the system 600. In this example, the process 700 is performed using the controller 330 and one or more playback devices 310. It will be appreciated however, that in other examples, at least some of the functions associated with the controller 330 may be performed by another portable device, such as a portable playback device, for example. In some examples, certain aspects of the process 700 can be performed at the controller 330 and / or at one or more playback devices 310 in the MPS 100. Examples of the process 700 can be distributed across multiple devices, as indicated in FIG. 7. In several embodiments, a coordinator device is designated to collect and store signal information from other playback devices and / or from the controller 330. Coordinator devices can be selected from the available devices of the MPS 100 based on one or more of several factors, including (but not limited to) RSSI of signals received from the controller 330 at the playback devices 310, frequency of use, device specifications (e.g., number of processor cores, processor clock speed, processor cache size, non-volatile memory size, volatile memory size, etc.). For example, a particular playback device 310 can be selected as a coordinator device based on how long its processor has been idle, so as not to interfere with the operation of any other devices during playback (e.g., selecting the playback device 310c located in the second bedroom 301d that is used infrequently). In certain embodiments, the coordinator device may be the controller 330.

[0148] At operation 702, a user may interact with the controller 330 via the user interface 133, for example, indicating that the user intends to select a target device for some activity (e.g., to start playback of audio, to form or break a bonded group, to change one or more characteristics of an ongoing playback session, such as to swap the playback session to another device, to change the volume of the audio, or to change the audio content being played, etc.).

[0149] Accordingly, at operation 704, the controller 330 may direct the coordinator device to initiate a beaconing session. Also at operation 704, the controller 330 may display initial information to the user on the user interface (e.g., a “home” screen for target device selection, and / or other information).

[0150] At operation 706, the coordinator device initiates the beaconing session. In other examples, the controller 330 may initiate the beaconing session, rather than instructing the coordinator device to do so. In some examples, initiating the beaconing session includes broadcasting, by the coordinator device or the controller 330, a wireless signal containing a beaconing instruction. Based on detecting the wireless signal, participating devices engage in a beaconing session 708. Participating devices may be all or some of the playback devices 310 in the environment 300. In some examples, participating devices may include all the playback devices 310 that detect the wireless signal containing the beaconing instruction and have the capability to transmit beacon signals 304. For example, at the time of any given beaconing session 708, some playback devices 310 may be powered off and may not detect the wireless signal. In some examples, based on one or more factors including the signaling technology / protocol used for beaconing, one or more of the playback devices 310 may detect the wireless signal, but be unable to participate in the beaconing session.

[0151] During the beaconing session 708, each of the participating devices broadcasts one or more beacon signals 304 at operation 710. According to certain examples, the beacon signals 304 are BLE signals. Accordingly, the participating devices may each include a wireless communication interface that includes a BLE radio, as described above. In some examples, the system is configured to employ passive reception in which, during the beaconing session 708, the controller 330 collects, at operation 712, beacon signals 304 emitted by the participating nodes, as described above. In other examples, as also described above, the participating nodes can also collect, during the beaconing session 708 and at operation 714, beacon signals 304 emitted by the controller 330. Accordingly, in some examples, at operation 712, the controller 330 also transmits one or more beacon signals 304. In such examples, each participating device, and the controller 330, is able to transmit and receive the beacon signals 304. Similar to the participating devices, the controller 330 may also include a wireless communication interfacethat allows for reception and optionally also transmission of BLE beacon signals. In some examples, the wireless signal transmitted to initiate the beaconing session 708 includes timing / synchronization information such that all the participating playback devices 310 and the controller 330 conduct the beaconing session during substantially the same overlapping time window.

[0152] Operations 710 (transmitting beacon signals 304) and 714 (detecting beacon signals 304) may be performed together during the beaconing session 708. In some examples, signaling during the beaconing session 708 is accomplished with standard HCI commands from the BLUETOOTH 5.3 core specification. However, in other examples, other signaling methodologies can be used. In examples, the signaling approach used during the beaconing session 708 does not require meticulous scheduling of the individual transmissions of beacon signals 304 from the participating devices, but rather just an alignment of the overall time window corresponding to the beaconing session 708. This is because BLE transmitters are capable of switching between transmit and receive modes quickly and apply small random offsets (e.g., 0 - 10ms) to each scheduled transmit time. Furthermore, the beacon signals 304 can be made to be very short transmissions. Accordingly, the random variation and short transmission time can be leveraged to avoid signal collisions during the beaconing session 708.

[0153] In some examples, the time duration of the beaconing session 708 is selected to allow sufficient time to perform signal measurements based on the beacon signals 304 collected at each participating device and the controller 330 to produce the input data 604. As discussed above, in some examples, the signal measurement data includes signal statistics obtained from a plurality of beacon signals collected at the controller 330 and / or each participating device. Accordingly, in such examples, the time duration of the beaconing session 708 should be long enough to allow the controller 330, and optionally each participating device, to collect a sufficient number of beacon signals to acquire statistically relevant data. In some examples, the time duration of the beaconing session 708 is approximately one second. The time duration of the beaconing session 708 may be selected based on a combination of factors, including, for example, (i) the time needed to collect a sufficient number of sample measurements, as discussed above; (ii) minimizing latency in the user interaction (e.g., the time between when the user begins interaction with the controller at operation 702 and the time at which the system can provide a recommendation for target device selection); and (iii) minimizing the time that the listening / collection event(s) occurring at operations 712 and / or 714 may impede device performance when the BLE antenna is shared with other components / functionality (e.g., witha WI-FI radio). Some playback devices 310 may include a dedicated BLE antenna, and therefore, factor (iii) may not be a consideration for such playback devices.

[0154] In some examples, BLE beacon signals may be transmitted at a rate of approximately 10 Hz. This rate is fast enough to have a robust signal measurement from the participating nodes for target device determination. In examples, an approximately 10 Hz sample rate with a one second beaconing time window allows for approximately 10 transmissions from the controller 330 and a similar number from each of the participating devices, which can then be detected by the controller 330.

[0155] As described above, during the beaconing session 708, the controller 330 (at operation 712) and / or the participating nodes (at operation 714) can acquire various signal measurements based on the collected beacon signals at each individual device. In some examples, the signal measurements include the RS SI value of each received beacon signal 304, the median signal strength (or RSSI value) for the group of beacon signals 304 collected during the beaconing session 708, and the standard deviation of the signal strength (e.g., RSSI values) over the group of collected beacon signals. The signal measurements may further include a count of the total number of beacon signals detected at a playback device 310 or controller 330 during the beaconing session 708.

[0156] FIG. 8 illustrates an example of a BLE signal distribution from one playback device 310 during the beaconing session 708. As described above, in certain examples, each BLE beacon signal 304 includes an advertisement packet to read its RSSI value. Thus, the “listening” device (whether the controller 330 or a participating device) may obtain the RSSI values for each detected beacon signal 304. In addition, the listening device may determine the median signal strength, or median RSSI value 802, for the group / set of detected beacon signals and the standard deviation 804 of the signal strength. These statistical values 802, 804 may be used as input features for the model 608. Another feature that can be determined and extracted from the signal data is a count of the total number of beacon signals 304 detected at the listening device during the beaconing session 708, as described above. This information may be determined by each listening device in the network. Thus, the total number of features that can be included in the input data 604 for a given beaconing session scales with the number of participating nodes. Additional features may include information such as the time of day of the beaconing session and / or recency of use of each participating playback device 310.

[0157] Referring again to FIG. 7, at operation 716, a computation device collects reporting signals from the controller 330 and each of the participating nodes to acquire the input data 604 that will be used by the MPC 602 to produce a target device prediction. The MPC 602 maythus be operated on the computation device. In some examples, the computation device is the same playback device 310 that acts as the coordinator device that initiates the beaconing session at operation 706. In other examples, the coordination device and the computation device can be different playback devices 310. In other examples, the controller 330 can be the computation device. The reporting signals may include the determined signal statistics discussed above (e.g., median RSSI value, standard deviation of signal strength, and count of detected beacon signals) and optionally the read RSSI values of each detected beacon signal 304. In addition, the reporting signals may each include identification information as to the playback device 310 (or controller 330) that is providing the reporting signal. In the case of the controller 330, which potentially receives beacon signals 304 from multiple playback devices 310, the reporting signal from the controller 330 may include playback device identity information corresponding to each individual beacon signal 304 or to sets of beacon signals received from the same playback device 310. As discussed above, each beacon signal 304 may include a transmission identifier and / or may comprise an individual sequence of tones or some other distinguishing characteristic that identifies the source of that beacon signal. Accordingly, this identification information may be included in the reporting signals provided from the controller 330 to the computation device at operation 716.

[0158] Also at operation 716, the computation device assembles, from the reporting signals and any signal measurement data acquired during the beaconing session 708 by the computation device itself, the input data 604 that is supplied to the MPC 602. The computation device thus can collect and store the input data 604. In some examples, the input data is stored in a matrix or other data structure in a memory (e.g., memory 112b) or other computer / machine readable storage device that is part of the computation device.

[0159] As described above, the signal measurement data collected from each participating device (and the controller 330) can be combined by the computation device to produce a signal pattern that corresponds the beaconing session (e.g., at operation 504 discussed above), and therefore to a particular location of the controller 330. For example, referring to FIG. 9A, there is illustrated an example of the environment 300 including playback devices 3 lOa-d, which are participating nodes in this example, and the controller 330 located at a position 1. A corresponding signal pattern 902 is illustrated in FIG. 9B. In this example, the signal pattern 902 includes signal data sets 904a, 904b, 904c, and 904d corresponding to the beacon signals detected from each of the four participating devices, namely, playback devices 310a-d in this example. The signal data sets 904a-d also include the median signal strength (Mi), the spread, or standard deviation, Si, and the count, Ci, of the total number of beacon signals collected bythe respective device during the beaconing session 708 (i = a, b, c, d). Thus, this signal pattern 902 may comprise the input data 604 that is provided to the MPC 602, and features extracted by the data sampler 610 may include the statistical information, Mi, Si, Ci, optionally along with other feature data as described above.

[0160] In the example of FIG. 9B, it can be seen that there is considerable overlap in the RSSI values of the beacon signals received from (and optionally at) the playback device 310c and those received from (and optionally at) the playback device 310d. However, the RSSI values collectively are much higher for the playback device 310a, which in the example of FIG. 9 A, is closest to the controller 330. Thus, in this particular example, the RSSI values may indicate that playback device 310a is a likely target device. However, as discussed above, in other examples, the RSSI values may indicate more ambiguity with respect to the potential target device.

[0161] As described above, in certain examples, the beacon signals 304 can also be transmitted by the controller 330 and detected by the playback devices 3 lOa-d. Accordingly, each playback device 310a-d may produce signal data based on the set of beacon signals emitted by the controller 330 that it detects. Accordingly, in some examples, the signal data sets 904a-d may be based on a combination of the signal data accumulated from beacon signals emitted by the respective playback device 310a-d and detected at the controller 330 and signal data accumulated from beacon signals emitted by the controller 330 and detected at the same respective playback device 310a-d. Such two-way signal exchange and corresponding combination of the signal data may add robustness to the signal pattern generation process.

[0162] According to certain examples, the beacon signals 304, or the collected RSSI values, can be pre-processed at either operation 712 / 714 or operation 716 to be scaled appropriately. As described above, in some examples a reference pattern produced based on the reference signals 306 can be used to normalize / scale the beacon signals. In other examples, a min / max scaling operation can be applied at operation 716 to all features in the data set.

[0163] In some examples, each reporting signal sent to the computation device at operation 716 only includes the three features discussed above, namely the median RSSI value, the standard deviation, and the count. Thus, transmitting the reporting signals to the computation device can occupy very little time and add very little latency to the process. Accordingly, the bandwidth of the BLE radios may not need to support the size of a full dataset, rather just the small content of the reporting signals. In some examples, the reporting signals are not sent via BLE, but are instead transmitted via a WI-FI channel in a network that communicatively couples the playback devices 310 and the controller 330 (e.g., the network 104 of FIG. IB). Byusing distribution data (e.g., the statistical calculations discussed above), rather than raw measurement data, the data set is further compressed, which may be advantageous both in terms of transmission and in terms of the time used by the MPC 602 to process the input data 604 and run the model 608 based on the input data 604.

[0164] At operation 718, the computation device operates the MPC 602 to run the model 608 based on the input data 604 to produce a predicted target device. In some examples, the model 608 is a logistic regression model with 5 fold cross validation; however, in other examples, other model kernels can be used. A logistic regression model may offer advantages due to its low computational complexity and ability to be trained on a relatively small training data set. In contrast, a neural network, for example, may require a large training data set and comparatively more computational resources to provide predictions with reasonable accuracy. In addition, logistic regression models are low latency models. As discussed above, the time duration of the beaconing session can be selected based on various factors, including the amount of time needed to obtain a statistically relevant data set (e.g., to collect at each listening device, a sufficient number of beacon signals 304 to produce a distribution that has a meaningful median and standard deviation). In some environments, the signal paths may be subject to large degrees of attenuation (due to distance or obstruction), and as a result, fewer beacon signals may be detected by the listening devices. Such factors like these affect the ability to accurately measure the signal strength distribution, therefore, it may be advantageous to select a sufficiently long time window for the beaconing session 708 that allows a reasonable modeling of the measurement, while also considering the influence on latency in providing a prediction to the user. Furthermore, for a logistic regression model 608, the prediction accuracy, and the error on this accuracy is mostly continuously increasing over the span of potential signal collection windows, with a rapid increase of an additional 10% in accuracy up to a 1 second window. Accordingly, model prediction accuracy may be a significant factor in selecting the time duration of the beaconing session.

[0165] In order to select an appropriate time window compared to prediction accuracy, the influence on a user’s experience can be considered, knowing that an incorrect prediction will likely result in a user being dissatisfied and leave the user confused and / or annoyed. This may be a stronger influence on the user experience over small increases in latency, which may otherwise be hidden by user interactions. In some examples, the system can be configured with a 98% accuracy set by the decision threshold 616 and a 1 second time window for the beaconing session 708. However, various other combinations can be selected in other examples, as will be appreciated given the benefit of this disclosure.

[0166] At operation 720, provided that the target device prediction produced by the model 608 exceeds the uncertainty threshold 618 (and optionally the decision threshold 616), a predicted target device can be offered / suggested to the user via the user interface 133 of the controller 330. Referring to FIG. 10, in some examples, this may be accomplished by displaying the predicted target device first in a list 1002 of available playback devices 310 that could be potential target devices, and / or by highlighting or otherwise emphasizing the predicted target device such that the user may be drawn to that selection. As discussed above, in the example of FIG. 9B, the signal pattern 902 may indicate that the playback device 310a located in the office 301c may be the correct target device. Accordingly, FIG. 10 shows an example of the display list 1002 corresponding to this example.

[0167] At operation 722, the user may, through the user interface 133, confirm or reject the predicted target device offered by the system. As described above, this passive user feedback can be used to label the data. For example, if the user selects the suggested target device, a “correct” label can be applied, whereas if the user selects a different playback device from the list (indicating that the prediction was incorrect), an “incorrect” label can be applied. In some examples, the screen displaying the list 1002 with the highlighted suggested target device may “time out” after a certain period of user inactivity. In such instances, the system can interpret this as confirmation of the selection by the user, and the predicted target device can be selected, and a correct label applied. In this manner, signal patterns for various locations of the controller 330 can be linked to particular target devices, and the labeled data can be used to re-train the model 608 (continual learning, as described above) such that the model learns to predict target devices for similar signal patterns with greater and greater accuracy (higher associated probabilities). An example is illustrated in FIGS. 11 A and 1 IB.

[0168] FIG. 11 A shows an example of the environment 300 with four locations (1 to 4) of the controller 330. FIG. 11B shows corresponding signal patterns 902 for each of the four locations. In this example, the signal patterns 902 associated with the first three locations (1, 2, 3) have been labeled with the correct corresponding target device (labels 1102), in this example, playback device 310a.

[0169] At operation 724, such labeled signal patterns 902 can be stored by the computation device.

[0170] In certain examples, operation 718 described above utilizes an initial model 608, indicated at 726. The initial model therefore lacks the benefit of the user feedback provided at operation 722. Accordingly, at operation 728, the initial model can be trained using the labeled data acquired and stored at operation 724 to produce an updated model, indicated at 730, thatcan be used for the next iteration of the process 700. In this manner, the system can apply continued learning to gather new training data over time and improve prediction accuracy as the amount of local labeled input data increases with user interactions with the MPS 100.

[0171] Thus, referring again to FIGS. 11 A and 1 IB, if during an iteration of the process 700, the signal pattern 902 corresponding to location 4 is obtained, based on similarity with the labeled signal patterns 902 for locations 1-3, the MPC 602 may predict that the playback device 310a is again the correct target device for interaction. Through numerous iterations of the process 700, labeled signal patterns for multiple locations throughout the environment 300 may be obtained and stored at operation 724. An example of sixteen different locations, and corresponding labeled signal patterns 902, is illustrated in FIGS. 12A and 12B.

[0172] According to certain examples, the sophistication or accuracy of the initial model applied indicated at 726 may depend, in part, on the number of times the process 700 has been performed previously within the environment 300. For example, as described above, the model 608 may continue to be trained and updated over time, incorporating passive user feedback to provide the labels 1102, and as such, the amount of labeled local data available may increase over time. Accordingly, the confidence metric (certainty) associated with the prediction may also increase over time. As described above, these factors may influence the manner in which the MPC 602 provides (or does not provide) recommendations to the user, and / or whether the MPC 602 automatically directs the predicted target device to be selected, rather than merely offering a suggestion to the user. In some instances, in early stages when little labeled local data is available and / or the prediction uncertainty is fairly high, the controller 330 can be configured to seek active user feedback by asking the user (e.g., via a visual message or voice query) whether the predicted target device is the playback device the user intends to interact with, and providing an easy path by which the user can select the correct target device (e.g., a clickable list) if the predicted target device is incorrect. This interactive dynamic that can be provided via the user interface 133 may allow the user to label the data and may translate into providing a personalized user experience that goes beyond relying on, or attempting to perceive the presence or absence of, physical boundaries, such as walls. Through continued usage of the system 600, users are more likely to be presented with the correct target device that they intend to interact with.

[0173] As also described above, in certain examples, the MPC 602 can be configured such that a recommended target device is only offered to the user if the uncertainty associated with the prediction falls within the “trust region” defined by the confidence element 614. Similarly, the level of intrusiveness (e.g., from no action, to a recommendation, to automatic selection) withwhich the MPC 602 presents a recommended target device may vary based on the confidence metric and the user preferences 612, as described above. Thus, the system can be configured to provide an adaptive, user-driven experience that can be tailored to individual users and / or environments and accommodate changes, in the environment and / or user routines or preferences, over time.

[0174] As described above, in certain examples, in addition to producing the signal patterns 902 based on the beacon signals 304, the system can be configured to generate reference patterns based on the reference signals 306. FIG. 12A shows an example of reference signals 306 transmitted between the playback devices 310a-d. As described above, the reference pattern can be used for various purposes, including providing a baseline or reference framework for comparing the signal patterns 902. In some examples, the reference pattern can be used to “localize” the signal patterns 902, such that the signal patterns can be linked to particular relative locations within the environment 300. In other words, the reference pattern can be used to identify the locations (e.g., locations 1-16 of FIG. 12B) in terms of relative proximity to one or more of the playback devices 310 in the environment 300.

[0175] As described above, in certain examples, contextual location information, such as room detection, can also (or alternatively) be applied to identify the locations in terms of relative positioning within the environment 300. For example, acoustic or ultrasonic signaling can be used for presence / room detection, as described above with reference to FIG. 2. This information, optionally in combination with the reference pattern, can be used to localize the locations, and therefore the associated signal patterns. Accordingly, the signal patterns 902, and their associated activities, can be tied to specific relative locations within the environment. This can allow for automatic location-based behavior to be implemented. For example, as described above, if a portable playback device is moved into a room containing a home theater set-up, and therefore may be a target device for bonding with one or more other home theater devices, the system can be configured to perform an automatic centering activity to prepare the portable playback device for becoming a home theater satellite.

[0176] In addition, in some instances, changes in the reference pattern can indicate a significant change in the environment 300, which may impact the validity or usefulness of previously stored labeled data. For example, FIG. 12C illustrates an example in which the playback device 310a has been moved from the office 301c into the master bedroom 301a. As a result, and as may be seen by comparing FIG. 12A to FIG. 12C, the reference pattern produced from the reference signals 306 may be substantially different. Also as a result, it can be seen with reference to FIGS. 12A-C, that the labeled “correct” target device 310a associated with thesignal patterns 902 for locations 1-4 may no longer be valid. Similarly, since the playback device 310a is now in the same room as the playback device 310b, the labels of playback device 310b as the “correct” target device for associated with the signal patterns 902 for locations 5- 9 may also no longer be valid.

[0177] Accordingly, in some examples, recognition of a change in the reference pattern can trigger the system 600 to discard some or all previously stored labeled data and cause the model 608 to undergo re-training to adapt to the new configuration of the environment 300. In certain examples, the system 600 can be configured to evaluate reference patterns according to a “change threshold,” such that if the difference between one instance of the reference pattern and another exceeds the change threshold, the system 600 may trigger retraining. This may allow the system to account for small variations in the reference pattern that may naturally occur due to differences in environmental conditions, even when no playback devices 310 have in fact moved or moved significantly. As discussed above, some of the playback devices 310 may be portable devices that are frequently moved, whereas others are stationary devices that rarely move. Accordingly, in some examples, one or more stationary devices can be designated as “anchor” devices and an anchor reference pattern can be established based on only those anchor devices. Accordingly, if the anchor reference pattern changes, this may trigger a complete retraining of the model 608 as described above. In other examples, a portable playback device may have contributed to the reference pattern, but may not be an anchor device, such that its movement alters the reference pattern but not the anchor reference pattern. In such cases, changes in the reference pattern, due to movement of the portable playback device, may trigger the system 600 to re-train the model 608 only with respect to labeled data that involved the portable playback device. In this manner, the system may seamlessly adapt to changes in the environment 300 while minimizing the likelihood that such changes result in incorrect target device predictions or selections that could cause a negative user experience.V. Conclusion

[0178] The above discussions relating to playback devices, controller devices, playback zone configurations, and media content sources provide only some examples of operating environments within which functions and methods described below may be implemented. Other operating environments and configurations of media playback systems, playback devices, and network devices not explicitly described herein may also be applicable and suitable for implementation of the functions and methods.

[0179] The description above discloses, among other things, various example systems, methods, apparatus, and articles of manufacture including, among other components, firmwareand / or software executed on hardware. It is understood that such examples are merely illustrative and should not be considered as limiting. For example, it is contemplated that any or all of the firmware, hardware, and / or software aspects or components can be embodied exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Accordingly, the examples provided are not the only ways to implement such systems, methods, apparatus, and / or articles of manufacture.

[0180] Additionally, references herein to “embodiment” means that a particular element, structure, or characteristic described in connection with the embodiment can be included in at least one example embodiment disclosed herein. The appearances of this phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments mutually exclusive of other embodiments. As such, the embodiments described herein, explicitly and implicitly understood by one skilled in the art, can be combined with other embodiments.

[0181] The specification is presented largely in terms of illustrative environments, systems, procedures, steps, logic blocks, processing, and other symbolic representations that directly or indirectly resemble the operations of data processing devices coupled to networks. These process descriptions and representations are typically used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art. Numerous specific details are set forth to provide a thorough understanding of the present disclosure. However, it is understood to those skilled in the art that certain embodiments of the present disclosure can be practiced without certain, specific details. In other instances, well known methods, procedures, components, and circuitry have not been described in detail to avoid unnecessarily obscuring aspects of the embodiments. Accordingly, the scope of the present disclosure is defined by the appended claims rather than the foregoing description of embodiments.

[0182] When any of the appended claims are read to cover a purely software and / or firmware implementation, at least one of the elements in at least one example is hereby expressly defined to include a tangible, non-transitory medium such as a memory, DVD, CD, Blu-ray, and so on, storing the software and / or firmware.VI. Additional Examples

[0183] The following examples pertain to further embodiments, from which numerous permutations and configurations will be apparent.

[0184] Example 1 provides a playback device comprising a wireless communication interface configured to support communication of data via at least one network protocol, at least oneprocessor; and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to, during a collection window, detect, via the wireless communication interface, one or more beacon signals emitted by a control device, determine, for the one or more beacon signals as a group, a median received signal strength indicator (RS SI) value and a standard deviation of a signal strength relative to a median RS SI value, to produce a first set of measured data, determine a first count of the one or more beacon signals detected during the collection window, detect, via the wireless communication interface, a plurality of reporting signals, each reporting signal emitted by a respective other playback device of a corresponding plurality of other playback devices, and each reporting signal including a second set of measured data representing (i) a median RS SI value and corresponding standard deviation value for a set of the beacon signals detected by the respective other playback device during the collection window, and (ii) a second count of the set of the beacon signals detected by the respective other playback device during the collection window, and based on the first set of measured data, the first count, and the plurality of reporting signals, identify from among a group including the playback device and the plurality of other playback devices, a proposed target playback device to receive, from the control device, playback instructions directing playback of audio content.

[0185] Example 2 includes the playback device of Example 1, wherein to identify the proposed target device, the at least one tangible non-transitory computer readable medium stores program instructions that are executable by the at least one processor to cause the playback device to apply a trained machine learning model to the first set of measured data, the first count, and the second set of measured data extracted from each of the plurality of reporting signals.

[0186] Example 3 includes the playback device of Example 2, wherein the machine learning model is configured to output a probability of a label based on a set of features, wherein the label is the proposed target playback device, and wherein the set of features includes the first set of measured data, the first count, and the plurality of second sets of measured data.

[0187] Example 4 includes the playback device of Example 3, wherein the trained machine learning model includes a logistic regression model.

[0188] Example 5 includes the playback device of any one of Examples 2-4, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to detect, via the wireless network interface, a message from the control device, the message including userfeedback regarding the proposed target playback device, and retrain the machine learning model based on the user feedback.

[0189] Example 6 includes the playback device of any one of Examples 1-5, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to transmit, via the wireless communication interface, a beacon signal.

[0190] Example 7 includes the playback device of Example 6, wherein at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to transmit the beacon signal based on an event signal received from the control device.

[0191] Example 8 includes the playback device of Example 7, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to transmit to the plurality of other playback devices, via the wireless network interface, an instruction to cause the plurality other playback devices to transmit the one or more beacon signals.

[0192] Example 9 includes the playback device of any one of Examples, 1-8, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to transmit to the control device, via the wireless network interface, an identification of the proposed target playback device.

[0193] Example 10 includes the playback device of any one of Examples 1-9, wherein the first set of measured data, the first count, and the plurality of second sets of measured data represent a beacon pattern corresponding to a location of the control device, and wherein the at least one tangible non-transitory computer readable medium stores program instructions that are executable by the at least one processor to cause the playback device to establish a reference pattern based on locations of the playback device and the plurality of other playback devices.

[0194] Example 11 includes the playback device of Example 10, wherein to establish the reference pattern, the at least one tangible non-transitory computer readable medium stores program instructions that are executable by the at least one processor to cause the playback device to transmit, via the wireless communication interface, a reference signal, detect, via the wireless communication interface, a plurality of other reference signals emitted by the plurality of other playback devices, detect, via the wireless communication interface, information indicative of a pattern of wireless signals between the plurality of other playback devices, andestablish the reference pattern based on detection of the plurality of other reference signals and the pattern of wireless signals between the plurality of other playback devices.

[0195] Example 12 includes the playback device of any one of Examples 1-11, wherein the at least one network protocol includes a BLUETOOTH LOW ENERGY protocol.

[0196] Example 13 includes the playback device of any one of Examples 1-12, wherein the collection window has a predetermined time duration.

[0197] Example 14 includes the playback device of Example 12, wherein the predetermined time duration is one second.

[0198] Example 15 includes the playback device of any one of Examples 1-14, wherein each of the one or more beacon signals includes a sequence of tones and a transmission identifier.

[0199] Example 16 includes the playback device of Example 15, wherein at least one of the sequence of tones or the transmission identifier identifies a playback device that is a source of the respective beacon signal.

[0200] Example 17 provides a media control device configured to control playback of audio content on a plurality of playback devices, the media control device comprising a user interface, a wireless communication interface configured to support communication of data via at least one network protocol, at least one processor, and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the media control device to detect, via the user interface, a user input indicative of an intent to initiate playback of the audio content on at least one of the plurality of playback devices, transmit, via the wireless communication interface, an instruction directing a first playback device of the plurality of playback devices to initiate a beaconing session among the plurality of playback devices, detect, via the wireless communication interface, one or more beacon signals emitted by one or more of the plurality of playback devices, for a group of the one or more beacon signals, determine a median received signal strength indicator (RSSI) value and a standard deviation of a signal strength relative to a median RSSI value to produce a set of measured data, determine a count of the one or more beacon signals detected during the beaconing session, transmit to the first playback device, via the wireless communication interface, a reporting signal including the set of measured data and the count, detect, via the wireless communication interface, a message from the first playback device, the message identifying a proposed target playback device for playback of the audio content, and display, via the user interface, a suggestion to the user to select the proposed target playback device for playback of the audio content.

[0201] Example 18 includes the media control device of Example 17, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the media control device to detect, via the user interface, user feedback regarding the proposed target playback device, and transmit, via the wireless communication interface, the user feedback to the first playback device.

[0202] Example 19 includes the media control device of Example 18, wherein the user feedback includes one of selection of the proposed target playback device for playback of the audio content or selection of another playback device from among the plurality of playback devices for playback of the audio content.

[0203] Example 20 includes the media control device of any one of Examples 17-19, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the media control device to during the beaconing session, transmit, via the wireless communication interface, a beacon signal.

[0204] Example 21 includes the media control device of any one of Examples 17-20, wherein the beaconing session has a predetermined time duration.

[0205] Example 22 includes the media control device of Example 21, wherein the predetermined time duration is approximately one second.

[0206] Example 23 includes the media control device of any one of Examples 17-22, wherein each respective beacon signal of the one or more beacon signals includes a sequence of tones and a transmission identifier.

[0207] Example 24 includes the media control device of Example 23, wherein at least one of the sequence of tones or the transmission identifier identifies a playback device, from among the plurality of playback devices, that transmitted the respective beacon signal.

[0208] Example 25 includes the media control device of any one of Examples 17-24, wherein the media control device is a mobile phone configured to run an audio content control application.

[0209] Example 26 provides a method comprising detecting an event indicative of a user intent to initiate an audio playback session on at least one playback device of a plurality of playback devices, based on detecting the event, initiating a beaconing session among the plurality of playback devices, during the beaconing session, collecting information indicative of a pattern of wireless signals between the plurality of playback devices, applying a trained machine learning model to the information to identify, from among the plurality of playback devices, aproposed target playback device for the playback session, and providing an identifier of the proposed target playback device.

[0210] Example 27 includes the method of Example 26, wherein initiating the beaconing session includes directing the plurality of playback devices to transmit beacon signals.

[0211] Example 28 includes the method of Example 27, wherein the beacon signals are BLUETOOTH LOW ENERGY signals.

[0212] Example 29 includes the method of one of Examples 27 or 28, wherein collecting the information includes detecting one or more beacon signals, for the one or more detected beacon signals as a group, determining a median received signal strength indicator (RSSI) value and a standard deviation of a signal strength relative to a median RSSI value, and determining a count of the one or more beacon signals detected during the beaconing session.

[0213] Example 30 includes the method of Example 29, wherein applying the trained machine learning model to the information comprises applying the model to the information to produce a label based on a set of features extracted from the information, wherein the label is the proposed target playback device, wherein the set of features includes the count and the RSSI values and standard deviations of the one or more beacon signals.

[0214] Example 31 includes the method of any one of Examples 26-30, wherein applying the trained machine learning model includes applying a logistic regression model.

[0215] Example 32 includes the method of any one of Examples 26-31, further comprising detecting user feedback regarding the proposed target playback device, and updating the machine learning model based on the user feedback.

[0216] Example 33 includes the method of any one of Examples 26-32, further comprising training the machine learning model based on a training dataset indicative of the pattern of wireless signals between the plurality of playback devices.

[0217] Example 34 provides a playback device comprising a wireless communication interface configured to support communication of data via at least one network protocol, at least one processor, and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to perform the method according to any one of Examples 26-33.

[0218] Example 35 provides a playback device comprising a wireless communication interface configured to support communication of data via at least one network protocol, at least one processor, and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to, during a beaconing session, transmit, via the wireless communication interface, a first beaconsignal, during the beaconing session, detect, via the wireless communication interface, one or more second beacon signals emitted by a control device, determine, for the one or more second beacon signals as a group, a median received signal strength indicator (RSSI) value and a standard deviation of signal strength relative to the median RSSI value, to produce a first set of measured data, determine a first count of the one or more second beacon signals detected during the collection window, detect, via the wireless communication interface, one or more reporting signals containing information indicative of patterns of beacon signals between the control device and a plurality of other playback devices, and based on the first set of measured data, the first count, and the information, identify from among a group including the playback device and the plurality of other playback devices, a proposed target playback device to receive, from the control device, playback instructions directing playback of audio content.

[0219] Example 36 includes the playback device of Example 35, wherein the information comprises a second set of measured data representing (i) a median RSSI value and corresponding standard deviation values for a set of the second beacon signals detected by a respective other playback device during the beaconing session, and (ii) a second count of the set of the second beacon signals detected by the respective other playback device during the beaconing session.

[0220] Example 37 includes the playback device of Example 36, wherein to identify the proposed target device, the at least one tangible non-transitory computer readable medium stores program instructions that are executable by the at least one processor to cause the playback device to apply a trained machine learning model to the first set of measured data, the first count, and the second set of measured data extracted from each of the one or more reporting signals.

[0221] Example 38 includes the playback device of Example 37, wherein the trained machine learning model includes a logistic regression model.

[0222] Example 39 includes the playback device of one of Examples 37 or 38, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to detect, via the wireless network interface, a message from the control device, the message including user feedback regarding the proposed target playback device, and retrain the machine learning model based on the user feedback.

[0223] Example 40 includes the playback device of any one of Examples 35-39, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device totransmit to the control device, via the wireless network interface, an identification of the proposed target playback device.

[0224] Example 41 includes the playback device of any one of Examples 35-40, wherein at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to transmit the first beacon signal based on an event signal received from the control device.

[0225] Example 42 includes the playback device of Example 41, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to initiate the beaconing session by transmitting to the plurality of other playback devices, via the wireless network interface, an instruction to cause the plurality other playback devices to transmit the second beacon signals.

[0226] Example 43 includes the playback device of any one of Examples 35-42, wherein the at least one network protocol includes a BLUETOOTH LOW ENERGY protocol.

[0227] Example 44 provides a media control device configured to control playback of audio content on a plurality of playback devices, the media control device comprising a user interface, a wireless communication interface configured to support communication of data via at least one network protocol, at least one processor, and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the media control device to detect, via the user interface, a user input indicative of an intent to initiate playback of the audio content on at least one of the plurality of playback devices, transmit, via the wireless communication interface, an instruction to initiate a beaconing session among the plurality of playback devices, detect, via the wireless communication interface, one or more beacon signals emitted by one or more of the plurality of playback devices, for the one or more beacon signals as a group, determine a median received signal strength indicator (RSSI) value and a standard deviation of a signal strength relative to the median RSSI value to produce a first set of measured data, determine a count of the one or more beacon signals detected during the beaconing session, detect, via the wireless communication interface, one or more reporting signals containing information indicative of patterns of beacon signals between the media control device and the plurality of playback devices, based on the first set of measured data, the first count, and the information, identify from among the plurality of playback devices, a proposed target playback device, and display, via the user interface, a suggestion to the user to select the proposed target playback device for playback of the audio content.

[0228] Example 45 includes the media control device of Example 44, wherein to detect the one or more reporting signals includes to detect, via the wireless communication interface, a plurality of reporting signals, each reporting signal emitted by a respective playback device of the plurality of playback devices, wherein the information contained in each reporting signal includes a second set of measured data representing (i) a median RS SI value and corresponding standard deviation value for a set of the beacon signals detected by the respective playback device during the beaconing session, and (ii) a second count of the set of the beacon signals detected by the respective playback device during the beaconing session.

[0229] Example 46 includes the media control device of one of Examples 44 or 45, wherein to identify the proposed target playback device, the at least one tangible non-transitory computer readable medium stores program instructions that are executable by the at least one processor to cause the media control device to apply a trained machine learning model to the first set of measured data, the first count, and the information.

[0230] Example 47 includes the media control device of Example 46, wherein the trained machine learning model includes a logistic regression model.

[0231] Example 48 includes the media control device of one of Examples 46 or 47, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the media control device to detect, via the user interface, user feedback regarding the proposed target playback device, and retrain the machine learning model based on the user feedback.

[0232] Example 49 includes the media control device of Example 48, wherein the user feedback includes one of selection of the proposed target playback device for playback of the audio content or selection of another playback device from among the plurality of playback devices for playback of the audio content.

[0233] Example 50 includes the media control device of any one of Examples 44-49, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the media control device to during the beaconing session, transmit, via the wireless communication interface, a beacon signal.

[0234] Example 51 includes the media control device of any one of Examples 44-50, wherein the beaconing session has a predetermined time duration.

[0235] Example 52 includes the media control device of Example 51, wherein the predetermined time duration is one second.

[0236] Example 53 includes the media control device of any one of Examples 44-52, wherein each respective beacon signal of the one or more beacon signals includes a sequence of tones and a transmission identifier.

[0237] Example 54 includes the media control device of Example 53, wherein at least one of the sequence of tones or the transmission identifier identifies a playback device, from among the plurality of playback devices, that transmitted the respective beacon signal.

[0238] Example 55 includes the media control device of any one of Examples 44-54, wherein the media control device is a mobile phone configured to run an audio content control application.

[0239] Example 56 provides a playback device comprising a wireless communication interface configured to support communication of data via at least one network protocol, at least one processor, and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to, during a beaconing session, detect, via the wireless communication interface, one or more beacon signals emitted by a control device, determine, for the one or more beacon signals as a group, a median received signal strength indicator (RS SI) value and a standard deviation of a signal strength relative to the median RSSI value, to produce a first set of measured data, determine a first count of the one or more beacon signals detected during the beaconing session, detect, via the wireless communication interface, a plurality of reporting signals, each reporting signal emitted by a respective other playback device of a corresponding plurality of other playback devices, and each reporting signal including a second set of measured data representing (i) a median RSSI value and corresponding standard deviation value for a set of the beacon signals detected by the respective other playback device during the beaconing session, and (ii) a second count of the set of the beacon signals detected by the respective other playback device during the beaconing session, and train a machine learning model based on the first set of measured data, the first count, and the second sets of measured data extracted from the reporting signals to identify from among a group including the playback device and the plurality of other playback devices, a proposed target playback device to receive, from the control device, playback instructions directing playback of audio content.

[0240] Example 57 includes the playback device of Example 56, wherein the at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to provide an identification of the proposed target playback device identified by the machine learning mode.

[0241] Example 58 includes the playback device of one of Examples 56 or 57, wherein the at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to initiate the beaconing session based on detecting an event indicative of a user intent to initiate an audio playback session on at least one playback device of the group.

[0242] Example 59 includes the playback device of any one of Examples 56-58, wherein the machine learning model includes a logistic regression model.

[0243] Example 60 includes the playback device of any one of Examples 56-59, wherein the at least one network protocol includes a BLUETOOTH LOW ENERGY protocol.

[0244] Example 61 provides a media playback system comprising one or more of the playback devices of any one of Examples 1-16, 34-43, or 56-60 and / or the media control devices of any one of Examples 17-25 or 44-55.

[0245] Example 62 provides a method comprising detecting, during a collection window, one or more beacon signals emitted by a control device, determining first statistical data for the one or more beacon signals as a group and a first count of the one or more beacon signals detected during the collection window, detecting a plurality of reporting signals emitted by respective playback devices of a plurality of playback devices, each reporting signal including second statistical data for a set of the beacon signals detected by each respective playback device during the collection window and a second count of the set of the beacon signals detected by the respective playback device during the collection window, and based on the first statistical data, the first count, and the plurality of reporting signals, identifying a proposed target playback device to receive, from the control device, playback instructions directing playback of audio content.

Claims

CLAIMS1. A playback device comprising: a wireless communication interface configured to support communication of data via at least one network protocol; at least one processor; and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to during a collection window, detect, via the wireless communication interface, one or more beacon signals emitted by a control device, determine, for the one or more beacon signals as a group, a median received signal strength indicator (RS SI) value and a standard deviation of a signal strength relative to a median RSSI value, to produce a first set of measured data, determine a first count of the one or more beacon signals detected during the collection window, detect, via the wireless communication interface, a plurality of reporting signals, each reporting signal emitted by a respective other playback device of a corresponding plurality of other playback devices, and each reporting signal including a second set of measured data representing (i) a median RSSI value and corresponding standard deviation value for a set of the beacon signals detected by the respective other playback device during the collection window, and (ii) a second count of the set of the beacon signals detected by the respective other playback device during the collection window, and based on the first set of measured data, the first count, and the plurality of reporting signals, identify from among a group including the playback device and the plurality of other playback devices, a proposed target playback device to receive, from the control device, playback instructions directing playback of audio content.

2. The playback device of claim 1, wherein to identify the proposed target device, the at least one tangible non-transitory computer readable medium stores program instructions that are executable by the at least one processor to cause the playback device to apply a trained machine learning model to the first set of measured data, the first count, and the second set of measured data extracted from each of the plurality of reporting signals.

3. The playback device of claim 2, wherein the machine learning model is configured to output a probability of a label based on a set of features; wherein the label is the proposed target playback device; and wherein the set of features includes the first set of measured data, the first count, and the plurality of second sets of measured data.

4. The playback device of claim 3, wherein the trained machine learning model includes a logistic regression model.

5. The playback device of any one of claims 2-4, wherein the at least one tangible non- transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to: detect, via the wireless network interface, a message from the control device, the message including user feedback regarding the proposed target playback device; and retrain the machine learning model based on the user feedback.

6. The playback device of any one of claims 1-5, wherein the at least one tangible non- transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to: transmit, via the wireless communication interface, a beacon signal.

7. The playback device of claim 6, wherein at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to: transmit the beacon signal based on an event signal received from the control device.

8. The playback device of claim 7, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to: transmit to the plurality of other playback devices, via the wireless network interface, an instruction to cause the plurality other playback devices to transmit the one or more beacon signals.

9. The playback device of any one of claims 1-8, wherein the at least one tangible non- transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to: transmit to the control device, via the wireless network interface, an identification of the proposed target playback device.

10. The playback device of any one of claims 1-9, wherein the first set of measured data, the first count, and the plurality of second sets of measured data represent a beacon pattern corresponding to a location of the control device; and wherein the at least one tangible non-transitory computer readable medium stores program instructions that are executable by the at least one processor to cause the playback device to: establish a reference pattern based on locations of the playback device and the plurality of other playback devices.

11. The playback device of claim 10, wherein to establish the reference pattern, the at least one tangible non-transitory computer readable medium stores program instructions that are executable by the at least one processor to cause the playback device to: transmit, via the wireless communication interface, a reference signal; detect, via the wireless communication interface, a plurality of other reference signals emitted by the plurality of other playback devices; detect, via the wireless communication interface, information indicative of a pattern of wireless signals between the plurality of other playback devices; and establish the reference pattern based on detection of the plurality of other reference signals and the pattern of wireless signals between the plurality of other playback devices.

12. The playback device of any one of claims 1-11, wherein the at least one network protocol includes a BLUETOOTH LOW ENERGY protocol.

13. The playback device of any one of claims 1-12, wherein the collection window has a predetermined time duration.

14. The playback device of any one of claims 1-13, wherein each of the one or more beacon signals includes a sequence of tones and a transmission identifier.

15. The playback device of claim 14, wherein at least one of the sequence of tones or the transmission identifier identifies a playback device that is a source of the respective beacon signal.

16. A media control device configured to control playback of audio content on a plurality of playback devices, the media control device comprising: a user interface; a wireless communication interface configured to support communication of data via at least one network protocol; at least one processor; and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the media control device to detect, via the user interface, a user input indicative of an intent to initiate playback of the audio content on at least one of the plurality of playback devices, transmit, via the wireless communication interface, an instruction directing a first playback device of the plurality of playback devices to initiate a beaconing session among the plurality of playback devices, detect, via the wireless communication interface, one or more beacon signals emitted by one or more of the plurality of playback devices, for a group of the one or more beacon signals, determine a median received signal strength indicator (RS SI) value and a standard deviation of a signal strength relative to a median RSSI value to produce a set of measured data, determine a count of the one or more beacon signals detected during the beaconing session, transmit to the first playback device, via the wireless communication interface, a reporting signal including the set of measured data and the count, detect, via the wireless communication interface, a message from the first playback device, the message identifying a proposed target playback device for playback of the audio content, and display, via the user interface, a suggestion to the user to select the proposed target playback device for playback of the audio content.

17. The media control device of claim 16, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the media control device to: detect, via the user interface, user feedback regarding the proposed target playback device; and transmit, via the wireless communication interface, the user feedback to the first playback device.

18. The media control device of claim 17, wherein the user feedback includes one of selection of the proposed target playback device for playback of the audio content or selection of another playback device from among the plurality of playback devices for playback of the audio content.

19. The media control device of one of any one of claims 16-18, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the media control device to: during the beaconing session, transmit, via the wireless communication interface, a beacon signal.

20. The media control device of any one of claims 16-19, wherein the beaconing session has a predetermined time duration.

21. The media control device of any one of claims 16-20, wherein each respective beacon signal of the one or more beacon signals includes a sequence of tones and a transmission identifier.

22. The media control device of claim 21, wherein at least one of the sequence of tones or the transmission identifier identifies a playback device, from among the plurality of playback devices, that transmitted the respective beacon signal.

23. The media control device of any one of claims 16-22, wherein the media control device is a mobile phone configured to run an audio content control application.

24. A method comprising: detecting an event indicative of a user intent to initiate an audio playback session on at least one playback device of a plurality of playback devices; based on detecting the event, initiating a beaconing session among the plurality of playback devices; during the beaconing session, collecting information indicative of a pattern of wireless signals between the plurality of playback devices, the wireless signals being BLUETOOTH LOW ENERGY signals; applying a trained machine learning model to the information to identify, from among the plurality of playback devices, a proposed target playback device for the playback session; and providing an identifier of the proposed target playback device.

25. The method of claim 24, wherein initiating the beaconing session includes directing the plurality of playback devices to transmit beacon signals.

26. The method of claim 25, wherein collecting the information includes: detecting one or more beacon signals; for the one or more detected beacon signals as a group, determining a median received signal strength indicator (RSSI) value and a standard deviation of a signal strength relative to a median RSSI value; and determining a count of the one or more beacon signals detected during the beaconing session.

27. The method of claim 26, wherein applying the trained machine learning model to the information comprises: applying the model to the information to produce a label based on a set of features extracted from the information; wherein the label is the proposed target playback device; wherein the set of features includes the count and the RSSI values and standard deviations of the one or more beacon signals.

28. The method of any one of claims 24-27, wherein applying the trained machine learning model includes applying a logistic regression model.

29. The method of any one of claims 24-28, further comprising: detecting user feedback regarding the proposed target playback device; and updating the machine learning model based on the user feedback.

30. The method of any one of claims 24-29, further comprising: training the machine learning model based on a training dataset indicative of the pattern of wireless signals between the plurality of playback devices.

31. A playback device comprising: a wireless communication interface configured to support communication of data via at least one network protocol; at least one processor; and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to perform the method according to any one of claims 24-30.

32. A playback device comprising: a wireless communication interface configured to support communication of data via at least one network protocol; at least one processor; and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to during a beaconing session, transmit, via the wireless communication interface, a first beacon signal, during the beaconing session, detect, via the wireless communication interface, one or more second beacon signals emitted by a control device, determine, for the one or more second beacon signals as a group, a median received signal strength indicator (RSSI) value and a standard deviation of signal strength relative to the median RSSI value, to produce a first set of measured data, determine a first count of the one or more second beacon signals detected during the collection window,detect, via the wireless communication interface, one or more reporting signals containing information indicative of patterns of beacon signals between the control device and a plurality of other playback devices, and based on the first set of measured data, the first count, and the information, identify from among a group including the playback device and the plurality of other playback devices, a proposed target playback device to receive, from the control device, playback instructions directing playback of audio content.

33. The playback device of claim 32, wherein the information comprises a second set of measured data representing (i) a median RS SI value and corresponding standard deviation values for a set of the second beacon signals detected by a respective other playback device during the beaconing session, and (ii) a second count of the set of the second beacon signals detected by the respective other playback device during the beaconing session.

34. The playback device of claim 33, wherein to identify the proposed target device, the at least one tangible non-transitory computer readable medium stores program instructions that are executable by the at least one processor to cause the playback device to apply a trained machine learning model to the first set of measured data, the first count, and the second set of measured data extracted from each of the one or more reporting signals.

35. The playback device of claim 34, wherein the trained machine learning model includes a logistic regression model.

36. The playback device of one of claims 34 or 35, wherein the at least one tangible non- transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to: detect, via the wireless network interface, a message from the control device, the message including user feedback regarding the proposed target playback device; and retrain the machine learning model based on the user feedback.

37. The playback device of any one of claims 32-36, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to:transmit to the control device, via the wireless network interface, an identification of the proposed target playback device.

38. The playback device of any one of claims 32-37, wherein at least one tangible non- transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to transmit the first beacon signal based on an event signal received from the control device.

39. The playback device of claim 38, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the playback device to: initiate the beaconing session by transmitting to the plurality of other playback devices, via the wireless network interface, an instruction to cause the plurality other playback devices to transmit the second beacon signals.

40. The playback device of any one of claims 32-39, wherein the at least one network protocol includes a BLUETOOTH LOW ENERGY protocol.

41. A media control device configured to control playback of audio content on a plurality of playback devices, the media control device comprising: a user interface; a wireless communication interface configured to support communication of data via at least one network protocol; at least one processor; and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the media control device to detect, via the user interface, a user input indicative of an intent to initiate playback of the audio content on at least one of the plurality of playback devices, transmit, via the wireless communication interface, an instruction to initiate a beaconing session among the plurality of playback devices, detect, via the wireless communication interface, one or more beacon signals emitted by one or more of the plurality of playback devices,for the one or more beacon signals as a group, determine a median received signal strength indicator (RS SI) value and a standard deviation of a signal strength relative to the median RS SI value to produce a first set of measured data, determine a count of the one or more beacon signals detected during the beaconing session, detect, via the wireless communication interface, one or more reporting signals containing information indicative of patterns of beacon signals between the media control device and the plurality of playback devices, based on the first set of measured data, the first count, and the information, identify from among the plurality of playback devices, a proposed target playback device, and display, via the user interface, a suggestion to the user to select the proposed target playback device for playback of the audio content.

42. The media control device of claim 41, wherein to detect the one or more reporting signals includes to detect, via the wireless communication interface, a plurality of reporting signals, each reporting signal emitted by a respective playback device of the plurality of playback devices, wherein the information contained in each reporting signal includes a second set of measured data representing (i) a median RS SI value and corresponding standard deviation value for a set of the beacon signals detected by the respective playback device during the beaconing session, and (ii) a second count of the set of the beacon signals detected by the respective playback device during the beaconing session.

43. The media device of one of claims 41 or 42, wherein to identify the proposed target playback device, the at least one tangible non-transitory computer readable medium stores program instructions that are executable by the at least one processor to cause the media control device to apply a trained machine learning model to the first set of measured data, the first count, and the information.

44. The media control device of claim 43, wherein the trained machine learning model includes a logistic regression model.

45. The media control device of one of claims 43 or 44, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the media control device to: detect, via the user interface, user feedback regarding the proposed target playback device; and retrain the machine learning model based on the user feedback.

46. The media control device of claim 45, wherein the user feedback includes one of selection of the proposed target playback device for playback of the audio content or selection of another playback device from among the plurality of playback devices for playback of the audio content.

47. The media control device of any one of claims 41-46, wherein the at least one tangible non-transitory computer readable medium further stores program instructions that are executable by the at least one processor to cause the media control device to: during the beaconing session, transmit, via the wireless communication interface, a beacon signal.

48. The media control device of any one of claims 41-47, wherein the beaconing session has a predetermined time duration.

49. The media control device of any one of claims 41-48, wherein each respective beacon signal of the one or more beacon signals includes a sequence of tones and a transmission identifier.

50. The media control device of claim 49, wherein at least one of the sequence of tones or the transmission identifier identifies a playback device, from among the plurality of playback devices, that transmitted the respective beacon signal.

51. The media control device of any one of claims 41-50, wherein the media control device is a mobile phone configured to run an audio content control application.

52. A playback device comprising: a wireless communication interface configured to support communication of data via at least one network protocol; at least one processor; and at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to during a beaconing session, detect, via the wireless communication interface, one or more beacon signals emitted by a control device, determine, for the one or more beacon signals as a group, a median received signal strength indicator (RS SI) value and a standard deviation of a signal strength relative to the median RS SI value, to produce a first set of measured data, determine a first count of the one or more beacon signals detected during the beaconing session, detect, via the wireless communication interface, a plurality of reporting signals, each reporting signal emitted by a respective other playback device of a corresponding plurality of other playback devices, and each reporting signal including a second set of measured data representing (i) a median RS SI value and corresponding standard deviation value for a set of the beacon signals detected by the respective other playback device during the beaconing session, and (ii) a second count of the set of the beacon signals detected by the respective other playback device during the beaconing session, and train a machine learning model based on the first set of measured data, the first count, and the second sets of measured data extracted from the reporting signals to identify from among a group including the playback device and the plurality of other playback devices, a proposed target playback device to receive, from the control device, playback instructions directing playback of audio content.

53. The playback device of claim 52, wherein the at least one tangible non-transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to: provide an identification of the proposed target playback device identified by the machine learning mode.

54. The playback device of one of claims 52 or 53, wherein the at least one tangible non- transitory computer readable medium storing program instructions that are executable by the at least one processor to cause the playback device to: initiate the beaconing session based on detecting an event indicative of a user intent to initiate an audio playback session on at least one playback device of the group.

55. The playback device of any one of claims 52-54, wherein the machine learning model includes a logistic regression model.