Generative audio playback using a wearable playback device

A wearable playback device dynamically generates and controls audio based on contextual data, addressing the limitations of existing systems by providing adaptive and immersive audio experiences.

JP2025534358APending Publication Date: 2025-10-15SONOS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025518632
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-22
Publication Date
2025-10-15

AI Technical Summary

Technical Problem

Existing media playback systems, such as those offered by SONOS, while allowing synchronized playback across multiple devices, do not support the dynamic generation and contextual control of audio content on wearable devices, limiting the ability to create unique and adaptive listening experiences.

Method used

A wearable playback device that detects user interactions and environmental data to automatically initiate and dynamically generate audio content based on contextual inputs, such as location, emotional state, and ambient conditions, enabling intelligent and adaptive audio playback.

Benefits of technology

Enables unique and adaptive audio experiences by generating and controlling playback based on real-time contextual data, providing personalized and immersive audio without requiring tethering to other devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025534358000001_ABST
    Figure 2025534358000001_ABST
Patent Text Reader

Abstract

Systems and methods are disclosed for playing generated media content through a wearable audio playback device, such as headphones. In one method, the wearable device detects that it is being worn by a user and obtains one or more input parameters via a network interface. Generated media content is then generated based on the one or more input parameters and played through the wearable playback device. In some examples, playback stops after the wearable playback device detects that it is no longer being worn by the user.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Application No. 63 / 377,776, filed September 30, 2022, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to consumer goods, and more particularly to methods, systems, products, features, services, and other elements of media playback systems directed to media playback or some aspect thereof. [Background technology]

[0003] Until SONOS, Inc. filed one of its early patent applications, titled "Method for Synchronizing Audio Playback between Multiple Networked Devices," in 2003 and began offering its first commercially available media playback system in 2005, options for accessing and listening to digital audio in outloud settings were limited. The SONOS Wireless HiFi System allows people to experience music from many sources through one or more networked playback devices. Through a software control application installed on a smartphone, tablet, or computer, users can play what they want in any room with a networked playback device. Furthermore, using the controller, multiple rooms can be grouped together for synchronized playback, with different songs streamed in each room with a playback device, or the same song can be played synchronously in all rooms.

[0004] With an ever-increasing interest in digital media, there is a continuing need to develop consumer-accessible technologies to further enhance the listening experience. [Brief explanation of the drawings]

[0005] The features, aspects, and advantages of the disclosed techniques may be better understood with regard to the following description, appended claims, and accompanying drawings, set forth below. [Figure 1A] 1 is a partial cutaway view of an environment having a media playback system configured in accordance with aspects of the disclosed technology. [Figure 1B] 1B is a schematic diagram of the media playback system and one or more networks of FIG. [Figure 1C] 1A and 1B are schematic diagrams of a media playback system including a wearable playback device for playing the generated audio content. [Figure 2A] Functional Block Diagram of an Exemplary Playback Device [Figure 2B] 2B is a schematic diagram of an example housing of the playback device of FIG. 2A. [Figure 2C] 2B is a schematic diagram of another example of a housing for the playback device of FIG. 2A. [Figure 2D] 2B is a schematic diagram of yet another example of a housing for the playback device of FIG. 2A; [Figure 2E] 2B is a schematic diagram of yet another example of a housing for the playback device of FIG. 2A; [Figure 3A] FIG. 1 illustrates an example configuration of a playback device according to an aspect of the present disclosure. [Figure 3B] FIG. 1 illustrates an example configuration of a playback device according to an aspect of the present disclosure. [Figure 3C] FIG. 1 illustrates an example configuration of a playback device according to an aspect of the present disclosure. [Figure 3D] FIG. 1 illustrates an example configuration of a playback device according to an aspect of the present disclosure. [Figure 3E] FIG. 1 illustrates an example configuration of a playback device according to an embodiment of the present disclosure. [Figure 4A] 1 is a functional block diagram of an exemplary control device according to aspects of the present disclosure. [Figure 4B] 1 is an illustration of a controller interface according to an aspect of the present disclosure; [Figure 4C] 1 is an illustration of a controller interface according to an aspect of the present disclosure; [Figure 5A] FIG. 1 illustrates an exemplary method for generative audio playback via a wearable audio playback device, according to aspects of the present disclosure. [Figure 5B] FIG. 1 illustrates an exemplary method for generative audio playback via a wearable audio playback device, according to aspects of the present disclosure. [Figure 5C] FIG. 1 illustrates an exemplary method for generative audio playback via a wearable audio playback device, according to aspects of the present disclosure. [Figure 5D] FIG. 1 illustrates an exemplary method for generative audio playback via a wearable audio playback device, according to aspects of the present disclosure. [Figure 6] FIG. 1 is a schematic diagram of a system for creating and playing generated media content in accordance with aspects of the present disclosure. [Figure 7] FIG. 1 is a schematic diagram of another system for creating and playing generated media content in accordance with aspects of the present disclosure. [Figure 8A] FIG. 1 illustrates an exemplary method for playing generated audio based on location, according to aspects of the present disclosure. [Figure 8B] FIG. 1 illustrates an exemplary method for playing generated audio based on location, according to aspects of the present disclosure. [Figure 8C] FIG. 1 illustrates an exemplary method for playing generated audio based on location, according to aspects of the present disclosure. [Figure 8D] FIG. 1 illustrates an exemplary method for playing generated audio based on location, according to aspects of the present disclosure. [Figure 9] FIG. 1 illustrates an example scenario involving playback of generated audio via a wearable playback device in a home environment, according to aspects of the present disclosure. [Figure 10A] FIG. 1 illustrates an exemplary method for exchanging playback of generated audio between wearable playback device(s) and out-loud playback device(s) according to aspects of the present disclosure. [Figure 10B]FIG. 1 illustrates an example scenario involving exchanging playback of generated audio between wearable playback device(s) and out-loud playback device(s) according to aspects of the present disclosure. [Figure 10C] FIG. 1 is a schematic diagram of a system for creating and playing generated media content in accordance with aspects of the present disclosure. [Figure 10D] FIG. 1 illustrates an example method for playing generated audio via both a wearable playback device and an out-loud playback device, according to aspects of the disclosure. [Figure 11] FIG. 1 illustrates an example rules engine for restricting playback of generated audio within an environment in accordance with aspects of the present disclosure.

[0006] Although the drawings are intended to illustrate exemplary embodiments, the present invention is not limited to the arrangements and apparatus shown in the drawings. In the drawings, identical reference numbers indicate at least generally similar elements. To facilitate the description of a particular element, the most significant digit(s) of a reference number will refer to the figure in which that element first appears. For example, element 103a is first introduced and described in FIG. 1A. DETAILED DESCRIPTION OF THE INVENTION

[0007] I. Overview Generative media content is content that is dynamically synthesized, created, and / or modified based on algorithms, whether implemented in software or physical models. Generative media content can change over time based solely on algorithms or in conjunction with contextual data (e.g., user sensor data, environmental sensor data, occurrence data). In various examples, such generative media content can include generative audio (e.g., music, ambient sounds, etc.), generative visual images (e.g., abstract visual designs that dynamically change shape, color, etc.), or any other suitable media content or combination thereof. As described elsewhere herein, generative audio can be generated, at least in part, via algorithms and / or non-human systems that utilize rule-based computations to generate novel audio content.

[0008] Because generated media content can change dynamically in real time, it enables unique user experiences not available using traditional media playback of pre-recorded content. For example, the generated audio can be endless and / or dynamic audio that changes as inputs to the algorithm (e.g., input parameters related to user input, sensor data, media source data, or any other suitable input data) change. In some examples, the generated audio can be used to steer a user's mood toward a desired emotional state, with one or more characteristics of the generated audio changing in response to real-time measurements that reflect the user's emotional state. As used in examples of the present technology, a system can provide generated audio based on the user's current and / or desired emotional state, based on the user's activity level, based on the number of users present in the environment, or based on any other suitable input parameters.

[0009] Listening to audio content, whether generated or pre-existing audio, on a wearable playback device (e.g., headphones, earphones) typically requires the wearable playback device to be tethered or connected to another device, such as a smartphone, tablet, or laptop. Certain wearable playback devices (e.g., WiFi-enabled devices) may not require tethering to another local device but may still require the user to affirmatively initiate audio playback (e.g., via a voice command). In some cases, it may be desirable for a listener to use a wearable device that is not connected to any other device and for playback to be initiated automatically, optionally responding to contextual data (e.g., data related to the listener's environment).

[0010] The present technology relates to a wearable playback device configured to play generated media content and intelligently control such playback based on contextual data. For example, the wearable playback device can detect a contact with at least one of a user's ears (e.g., by detecting headphones on one ear or by detecting earphones being placed on at least one of the user's ears) and automatically begin playing a soundscape (e.g., generated audio, media content, other sounds, such as sounds related to the user's environment). The generated audio soundscape can include a generative musical composition based at least in part on one or more media content stems and / or audio cues derived from contextual data. Contextual data can include information related to the user's environment (e.g., location, position or orientation relative to the environment, time of day, temperature, circadian rhythm, humidity, number of nearby people, ambient light level), or other types of indicators (e.g., doorbell, alarm, event). In some implementations, certain other indicators (e.g., doorbell, alarm, etc.) can be selectively passed through the wearable playback device to notify the user while they remain within the immersive audio experience. Additionally, audio captured in the environment can generally serve as input for generative media algorithms, and in some cases the output of the generative media algorithms can themselves serve as input for future iterations (i.e., resampling) of the generative media content.

[0011] In various examples, the contextual data can be obtained via on-board sensor(s) carried by the wearable playback device, sensor(s) associated with other playback devices in the environment, or any other suitable sensor data source. In particular examples, the wearable playback device can identify the user when the playback device is placed on the user's head, and the wearable playback device can further tailor the generated media content to the user's profile, current or desired emotional state, and / or other biometric data (e.g., brainwave activity, heart rate, breathing rate, skin hydration, ear shape, head direction or orientation). Furthermore, in some cases, playback of the generated audio content can be dynamically swapped or toggled between playback via the wearable playback device and alternative or simultaneous playback via one or more out-loud playback devices in the listening environment.

[0012] Although some embodiments described herein may refer to functions performed by certain actors, such as "users" and / or other entities, it should be understood that this description is for illustrative purposes only. The claims should not be construed as requiring actions by such example actors unless expressly required by the claim language itself.

[0013] II. Operating environment example 1A, 1B, and 1C illustrate an exemplary configuration of a media playback system 100 (or "MPS 100") in which one or more embodiments disclosed herein may be implemented. With reference to FIG. 1A, the illustrated MPS 100 is associated with an exemplary home environment having multiple rooms and spaces, sometimes collectively referred to as a "home environment," "smart home," or "environment 101." Environment 101 comprises a home with multiple rooms, spaces, and / or playback zones, including a master bathroom 101a, a master bedroom 101b (referred to herein as "Nick's Room"), a second bedroom 101c, a family room or study 101d, an office 101e, a living room 101f, a dining room 101g, a kitchen 101h, and an outdoor patio 101i. While particular embodiments and examples are described below in the context of a home environment, the techniques described herein may also be implemented in other types of environments. In some embodiments, for example, MPS 100 may be implemented in one or more commercial environments (e.g., a restaurant, mall, airport, hotel, retail store or other establishment), one or more vehicles (e.g., a sport utility vehicle, bus, car, ship, boat, airplane), multiple environments (e.g., a combination of home and vehicle environments), and / or another suitable environment where multi-zone audio may be desirable.

[0014] Within these rooms and spaces, the MPS 100 includes one or more computing devices. Referring jointly to FIGS. 1A, 1B, and 1C, such computing devices include playback devices 102 (individually identified as playback devices 102a-102p), network microphone devices 103 (individually identified as “NMDs” 103a-102i), and control devices 104a and 104b (collectively “control devices 104”). Referring to FIG. 1B, the home environment may include additional and / or other computing devices, including one or more smart lighting devices 108 (FIG. 1B), smart thermostat 110, and local network devices such as local computing device 105 (FIG. 1A). In the embodiments described below, one or more of the various playback devices 102 may be configured as portable playback devices, and other playback devices may be configured as stationary playback devices. For example, headphones 102o (FIG. 1B) may be a portable playback device, while bookshelf playback device 102d may be a stationary playback device. As another example, patio playback device 102c may be a battery-powered device, allowing it to be carried to various locations within or outside environment 101 when not plugged in, such as to an electrical outlet. Additionally, one or more of the various playback devices 102 may be configured as wearable playback devices (e.g., playback devices configured to be worn by, by, or about a user, such as headphones, earphones, or smart glasses with integrated audio transducers). Playback devices 102o and 102p (FIG. 1C) are examples of wearable playback devices. In contrast to wearable playback devices, one or more of the various playback devices 102 may be configured as out-loud playback devices (e.g., playback devices configured to output audio content for listeners and / or multiple users some distance from the playback device, as opposed to the private listening experience associated with a wearable playback device).

[0015] 1B , the various playback devices, network microphones, and control devices 102-104 of the MPS 100, and / or other network devices, may be coupled to one another via a local network 111, which may include a network router 109, via point-to-point connections, and / or via other connections, which may be wired and / or wireless. For example, playback device 102j in den 101d (FIG. 1A), which may be designated as the "left" device, may have a point-to-point connection with playback device 102a, also in den 101d, which may be designated as the "right" device. In a related embodiment, left playback device 102j can communicate with other network devices, such as playback device 102b, which may be designated as the "front" device, via point-to-point connections and / or other connections via the local network 111. The local network 111 may be, for example, a network interconnecting one or more devices within a limited area (e.g., a residence, an office building, an automobile, a personal workspace, etc.). The local network 111 may include, for example, one or more local area networks (LANs), such as a wireless local area network (WLAN) (e.g., a WI-FI network, a Z-WAVE network, etc.), and / or one or more personal area networks (PANs), such as a BLUETOOTH (registered trademark) network, a wireless USB network, a ZIGBEE (registered trademark) network, an IRDA network, etc.

[0016] 1B , MPS 100 can be coupled to one or more remote computing devices 106 via a wide area network (“WAN”) 107. In some embodiments, each remote computing device 106 can take the form of one or more cloud servers. Remote computing devices 106 can be configured to interact with computing devices within environment 101 in various ways. For example, remote computing devices 106 can be configured to facilitate streaming and / or playback control of media content, such as audio, in home environment 101.

[0017] In some implementations, the various playback devices, NMDs, and / or control devices 102-104 may be communicatively coupled to at least one remote computing device associated with a voice assistant service (“VAS”) and / or at least one remote computing device associated with a media content service (“MCS”). For example, in the example of FIG. 1B, remote computing device 106a is associated with VAS 190, and remote computing device 106b is associated with MCS 192. While the example of FIG. 1B shows only a single VAS 190 and a single MCS 192 for clarity, MPS 100 may be coupled to multiple different VASs and / or MCSs. In some implementations, the VAS may be operated by one or more of AMAZON, GOOGLE, APPLE, MICROSOFT, NUANCE, SONOS, or other voice assistant providers. In some implementations, the MCS is operated by one or more of SPOTIFY, PANDORA, AMAZON MUSIC, or other media content services.

[0018] 1B , the remote computing device 106 further includes a remote computing device 106c configured to perform certain operations, such as remotely facilitating media playback functions, managing device and system status information, and directing communications between devices of the MPS 100 and one or more VASs and / or MCSs. In one example, the remote computing device 106c provides a cloud server for one or more SONOS Wireless HiFi Systems.

[0019] In various implementations, one or more of the playback devices 102 may take the form of or include an on-board (e.g., integrated) network microphone device. For example, playback devices 102a-102e include or otherwise comprise corresponding NMDs 103a-103e, respectively. Playback devices that include or comprise an NMD may be referred to interchangeably herein as playback devices or NMDs, unless otherwise indicated in the description. In some cases, one or more of the NMDs 103 may be standalone devices. For example, NMDs 103f and 103g may be standalone devices. A standalone NMD may omit components and / or functionality typically included in a playback device, such as speakers and associated electronics. For example, in such cases, the standalone NMD may not generate audio output or may generate limited audio output (e.g., relatively low-quality audio output).

[0020] The various playback and network microphone devices 102 and 103 of the MPS 100 may each be given individual names, which may be assigned to each device by a user, such as during setup of one or more of these devices. For example, as shown in the example of FIG. 1B, a user may assign the name “Bookshelf” to playback device 102d because playback device 102d is physically located on a bookshelf. Similarly, NMD 103f may be assigned the name “Island” because it is physically located on an island countertop in kitchen 101h (FIG. 1A). For example, playback devices 102e, 102l, 102m, and 102n may be named “Bedroom,” “Dining Room,” “Living Room,” and “Office,” respectively. Additionally, particular playback devices may have functionally descriptive names. For example, playback devices 102a and 102b are assigned the names "Right" and "Front," respectively, because these two devices are configured to provide specific audio channels during media playback in a zone of den 101d (FIG. 1A). Patio playback device 102c may be named portable because it is battery-powered and / or easily portable to different areas of environment 101. Other naming conventions are possible.

[0021] An NMD can detect and process sounds from its environment, such as sounds containing background noise mixed with speech from a person in the vicinity of the NMD, as described above. For example, when sounds in the environment are detected by the NMD, the NMD can process the detected sounds to determine whether the sounds contain audio containing audio input intended for the NMD and, therefore, a particular VAS. For example, the NMD can identify whether the audio contains a wake word associated with a particular VAS.

[0022] In the depicted example of FIG. 1B, the NMD 103 is configured to interact with the VAS 190 via the local network 111 and / or router 109. Interaction with the VAS 190 may begin, for example, when the NMD identifies a possible wake word from detected sounds. This identification triggers a wake word event, causing the NMD to initiate transmission of detected sound data to the VAS 190. In some implementations, the various local network devices 102-105 (FIG. 1A) of the MPS 100 and / or the remote computing device 106c may exchange various feedback, information, instructions, and / or related data with a remote computing device associated with a selected VAS. Such exchanges may be related to or independent of a transmitted message including speech input. In some embodiments, the remote computing device(s) and media playback system 100 may exchange data via a communication path as described herein and / or using a metadata exchange channel as described in U.S. Patent Publication No. 2017-0242653, published August 24, 2017, and entitled "Voice Control of a Media Playback System," which is incorporated herein by reference in its entirety.

[0023] Upon receiving the stream of sound data, the VAS 190 determines whether speech input is present in the streamed data from the NMD, and if so, the VAS 190 also determines the intent contained in the speech input. The VAS 190 then sends a response back to the MPS 100, which may include sending a response directly to the NMD that generated the wake word event. The response is typically based on the intent that the VAS 190 determines is present in the speech input. As an example, in response to the VAS 190 receiving speech input with the utterance "Play Hey Jude by The Beatles," the VAS 190 may determine that the intent contained in the speech input is to start playback and further determine that the intent of the speech input is to play the particular song "Hey Jude." After these determinations, the VAS 190 may send a command to a particular MCS 192 to retrieve content (i.e., the song "Hey Jude"), which in turn provides (e.g., streams) this content to the MPS 100 directly or indirectly via the VAS 190. In some implementations, the VAS 190 can send a command to the MPS 100 that causes the MPS 100 to retrieve the content from the MCS 192 itself.

[0024] In certain implementations, NMDs can facilitate arbitration between each other when audio input is identified as sound detected by two or more NMDs located in close proximity to each other. For example, NMD-equipped playback device 102d in environment 101 (FIG. 1A) may be relatively close to NMD-equipped living room playback device 102m, and both devices 102d and 102m may potentially detect the same sound. In such cases, arbitration may be required regarding which device is ultimately responsible for providing detected sound data to the remote VAS. Examples of arbitration between NMDs are described, for example, in U.S. Patent Publication No. 2017-0242653, discussed above.

[0025] In certain implementations, an NMD may be assigned or otherwise associated with a designated or default playback device that may not include an NMD. For example, island NMD 103f in kitchen 101h (FIG. 1A) may be assigned to dining room playback device 102l, which is relatively close to island NMD 103f. In practice, the NMD may instruct its assigned playback device to play audio in response to a remote VAS receiving a voice input from the NMD for audio playback. This audio may have been sent by the NMD to the VAS in response to a user speaking a command to play a particular song, album, playlist, etc. Additional details regarding assigning NMDs and playback devices as designated or default devices are disclosed, for example, in the previously discussed U.S. Patent Publication No. 2017-0242653.

[0026] 1C , media playback system 100 can be configured to generate and play back generated media content via one or more wearable playback devices 102o, 102p and / or via one or more out-loud playback devices 102. In the illustrated example, wearable playback devices 102o, 102p each take the form of headphones, although any suitable wearable playback device can be used. Out-loud playback devices 102 can include any suitable device configured to output audio content for listening at loud volumes in an environment. Some out-loud playback devices include integrated audio transducers (e.g., sound bars, subwoofers), while others include amplifiers configured to provide an output signal that is played back through other devices (e.g., hub devices, set-top boxes, etc.).

[0027] 1C can be wired or wireless network connections, which may be facilitated at least in part through router 109. Wireless connections may include WiFi, Bluetooth, or any other suitable communication protocol. As shown, one or more local sources 150 may be connected to wearable playback device 102o and outloud playback device 102 via a network (e.g., via router 109). Local source(s) 150 may include any suitable media and / or audio source, such as a display device (e.g., a television, projector, etc.), a microphone, an analog playback device (e.g., a turntable), a portable data storage device (e.g., a USB stick), computer storage (e.g., a laptop computer hard drive), etc. These local source(s) 150 may optionally provide audio content that is played via wearable playback device 102o, 102p, and / or outloud playback device 102. In some examples, media content obtained from local sources 150 can provide input to the generated media module, such that the generated media content is based on and / or incorporates features of the media content from local sources 150.

[0028] The first wearable playback device 102o may be communicatively coupled to a control device 104 (e.g., a smartphone, tablet, laptop, etc.), for example, via WiFi, Bluetooth, or other suitable wireless connection. The control device 104 may be used to select content and / or otherwise control playback of audio via the wearable playback device 102o and / or any additional playback devices 102.

[0029] The first wearable playback device 102o is also optionally connected to a local area network via router 109 (e.g., via a WiFi connection). The second wearable playback device 102p can connect to the first wearable playback device 102o via a direct wireless connection (e.g., Bluetooth) and / or via a wireless network (e.g., a WiFi connection via router 109). The wearable playback devices 102o, 102p can communicate with each other to transfer audio content, timing information, sensor data, generated media content parameters (e.g., content models, algorithms, etc.), or any other appropriate information to facilitate the generation, selection, and / or playback of audio content. The second wearable playback device 102p can also optionally connect to one or more out-of-home playback devices via a wireless connection (e.g., a Bluetooth or WiFi connection to a hub device).

[0030] The media playback system 100 may also communicate with one or more remote computing devices 154 associated with media content providers and / or generated audio sources. These remote computing devices 154 may optionally provide media content, including generated media content. As described in more detail elsewhere herein, in some implementations, the generated media content may be generated via one or more generated media modules, which may be instantiated via the remote computing device 154, via one or more playback devices 102, and / or via some combination thereof.

[0031] The generated media module may generate the generated media based at least in part on input parameters, which may include sensor data (e.g., received from sensor data source(s) 152) and / or other suitable input parameters. With respect to sensor input parameters, the sensor data source(s) 152 may include data from any suitable sensors, regardless of the location of the various playback devices or the values ​​measured thereby. Examples of suitable sensor data include physiological sensor data obtained from biosensors, wearable sensors, and the like. Such data may include physiological parameters such as heart rate, respiration rate, blood pressure, brain waves, activity level, movement, and body temperature.

[0032] Suitable sensors include wearable sensors configured to be worn or carried by a user, such as a wearable playback device, headset, watch, mobile device, brain-machine interface, microphone, or other similar device. In some examples, the sensors are non-wearable sensors or are affixed to a fixed structure. The sensors can provide sensor data including, for example, data corresponding to brain activity, a user's mood or emotional state, voice, location, movement, heart rate, pulse, body temperature, and / or sweating.

[0033] In some examples, sensor data source(s) 152 include data obtained from networked device sensor data (e.g., Internet-of-Things (IoT) sensors, such as networked lights, cameras, temperature sensors, thermostats, presence detectors, microphones, etc.) Additionally or alternatively, sensor data source(s) 152 may include environmental sensors (e.g., measuring or indicating weather, temperature, time / day / week / month), as well as user calendar events, proximity to other electronic devices, historical usage patterns, blockchain or other distributed data sources, etc.

[0034] In one example, a user may wear a biometric device capable of measuring various biometric parameters, such as the user's heart rate and blood pressure. A generative media module (whether resident on the remote computing device 154 or on one or more local playback devices 102) may use these parameters to further adapt the generated audio, such as increasing the tempo of the music in response to detecting a high heart rate (which may indicate the user is engaged in high athletic activity) or decreasing the tempo of the music in response to detecting high blood pressure (which may indicate the user is stressed and may benefit from calming music). In yet another example, one or more microphones on the playback device may detect the user's voice. The captured audio data may be processed to, for example, determine the user's mood, age, and gender, identify a particular user among multiple users in a household, or determine other input parameters. Other examples are also contemplated.

[0035] Additional details regarding the generation and playback of generated media content can be found in International Patent Application Publication No. WO2022 / 109556, owned by the present applicant, entitled "Playback of Generated Media Content," the contents of which are incorporated herein by reference in their entirety.

[0036] Further aspects related to the different components of the exemplary MPS 100 and how the different components may interact to provide a media experience to a user can be found in the following sections. While the description herein may generally refer to the exemplary MPS 100, the techniques described herein are not limited to application within the home environment described above. For example, the techniques described herein are useful in other home environment configurations that include one or more or a small number of any of the playback devices, network microphones, and / or control devices 102-104. For example, the techniques herein can be utilized within an environment with a single playback device 102 and / or a single NMD 103. In some examples of such cases, the local network 111 (FIG. 1B) can be eliminated, and the single playback device 102 and / or single NMD 103 can communicate directly with the remote computing devices 106a-d. In some embodiments, a telecommunications network (e.g., an LTE network, a 5G network, etc.) can communicate with the various playback devices, network microphones, and / or control devices 102-104 independently of the local network 111.

[0037] Although a specific implementation of an MPS is described above with reference to Figures 1A-1C, there are numerous configurations of an MPS, including, but not limited to, those that do not interact with remote services, systems that do not include a controller, and / or other configurations appropriate to the requirements of a given application.

[0038] a. Playback & Network Microphone Device Example 2A is a functional block diagram illustrating one particular aspect of the playback device 102 of the MPS 100 of FIGS. 1A-1C. As shown, the playback device 102 includes various components, each of which is described in further detail below, and which may be operatively coupled to one another via a system bus, a communications network, or some other connection mechanism. In the illustrated example of FIG. 2A, the playback device 102 may be referred to as an "NMD-equipped" playback device because it includes components that support the functionality of an NMD, such as one of the NMDs 103 shown in FIG. 1A.

[0039] As shown, the playback device 102 includes at least one processor 212, which may be a clocked computing component configured to process input data according to instructions stored in a memory 213. The memory 213 may be a tangible, non-transitory, computer-readable medium configured to store instructions executable by the processor 212. For example, the memory 213 may be data storage that stores software code 214 executable by the processor 212 to implement certain functions.

[0040] In one example, these functions may include the playback device 102 obtaining audio data from audio sources that are other playback devices. In another example, the functions may include the playback device 102 transmitting audio data, detected sound data (e.g., corresponding to audio input), and / or other information to other devices on the network via at least one network interface 224. In yet another example, the functions may include the playback device 102 causing one or more other playback devices to play audio in synchronization with the playback device 102. In yet another example, the functions may include facilitating the playback device 102 pairing or otherwise bonding with one or more other playback devices to create a multi-channel audio environment. Many other example functions are possible, some of which are described below.

[0041] As noted above, certain features may include the playback device 102 synchronizing the playback of audio content with one or more other playback devices. During synchronized playback, a listener does not perceive differences in time delays between the playback of audio content by the synchronized playback devices. This is described in U.S. Patent No. 8,234,395, filed April 4, 2004, entitled "System and Method for Synchronizing Operation Among Multiple Independently Clocked Digital Data Processing Devices," the contents of which are incorporated herein by reference in their entirety.

[0042] To facilitate audio playback, the playback device 102 generally includes an audio processing component 216 configured to process the audio before the playback device 102 renders it. In this regard, the audio processing component 216 may include one or more digital-to-analog converters (“DACs”), one or more audio preprocessing components, one or more audio enhancement components, one or more digital signal processors (“DSPs”), etc. In some implementations, one or more of the audio processing components 216 may be subcomponents of the processor 212. In operation, the audio processing component 216 receives analog and / or digital audio and processes and / or otherwise intentionally alters the audio to generate an audio signal for playback.

[0043] The generated audio signals can be provided to one or more audio amplifiers 217 to be amplified and played through one or more speakers 218 operably coupled to the amplifiers 217. The audio amplifiers 217 can include components configured to amplify the audio signals to a level to drive the one or more speakers 218.

[0044] Each of the speakers 218 may include an individual transducer (e.g., a “driver”), or the speakers 218 may include a complete speaker system including an enclosure with one or more drivers. A particular driver of the speakers 218 may include, for example, a subwoofer (e.g., for low frequencies), a mid-range driver (e.g., for medium frequencies), and / or a tweeter (e.g., for high frequencies). In some cases, the transducers may be driven by individual corresponding audio amplifiers of the audio amplifiers 217. In some implementations, the playback device may not include the speakers 218 but may instead include a speaker interface for connecting the playback device to external speakers. In certain embodiments, the playback device does not include the speakers 218 or the audio amplifiers 217, but instead includes an audio interface (not shown) for connecting the playback device to an external audio amplifier or audio-visual receiver.

[0045] In addition to generating audio signals for playback by playback device 102, audio processing component 216 may be configured to process audio that is sent for playback to one or more other playback devices via network interface 224. In an example scenario, audio content to be processed and / or played by playback device 102 may be received from an external source (not shown) via an audio line-in interface (e.g., an auto-detecting 3.5 mm audio line-in connection) of playback device 102 or via network interface 224, as described below.

[0046] As shown, the at least one network interface 224 can take the form of one or more wireless interfaces 225 and / or one or more wired interfaces 226. The wireless interfaces enable the playback device 102 to communicate with other devices (e.g., other playback device(s), NMD(s), and / or control device(s)) according to a communication protocol (e.g., IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.11ad, 802.11af, 802.11ah, 802.11ai, 802.11aj, 802.11aq, 802.11ax, 802.11ay, 802.15, BLUETOOTH, 4G mobile communication standards, 5G mobile communication standards, etc.). The wired interface may enable the playback device 102 to communicate wirelessly with other devices (e.g., other playback devices, NMDs, and / or control devices). The wired interface may provide a network interface function for the playback device 102 to communicate with other devices via a wired connection according to a communication protocol (e.g., IEEE 802.3). Although the network interface 224 shown in FIG. 2A includes both a wired interface and a wireless interface, the playback device 102 may, in some embodiments, include only a wireless interface or only a wired interface.

[0047] Generally, the network interface 224 facilitates the flow of data between the playback device 102 and one or more other devices on a data network. For example, the playback device 102 may be configured to receive audio content over a data network from one or more other playback devices, network devices in a LAN, and / or audio content sources over a WAN, such as the Internet. In one example, audio content and other signals sent and received by the playback device 102 may be transmitted in the form of digital packet data consisting of an Internet Protocol (IP)-based source address and an IP-based destination address. In such cases, the network interface 224 may be configured to parse the digital packet data so that data destined for the playback device 102 is appropriately received and processed by the playback device 102.

[0048] 2A , the playback device 102 also includes an audio processing component 220 operably coupled to one or more microphones 222. The microphones 222 are configured to detect sound (i.e., acoustic waves) in the environment of the playback device 102, which is provided to the audio processing component 220. More specifically, each microphone 222 is configured to detect and convert sound into a digital or analog signal representative of the detected sound, thereby enabling the audio processing component 220 to perform various functions based on the detected sound, as described in further detail below. In one implementation, the microphones 222 are arranged as an array of microphones (e.g., an array of six microphones). In some embodiments, the playback device 102 includes six or more microphones (e.g., eight microphones or twelve microphones) or fewer than six microphones (e.g., four microphones, two microphones, or a single microphone).

[0049] In operation, the audio processing component 220 is generally configured to detect and process sounds received via the microphone 222, identify potential voice inputs among the detected sounds, and extract the detected sound data so that a VAS, such as the VAS 190 (FIG. 1B), can process the identified voice inputs among the detected sound data. The audio processing component 220 includes one or more analog-to-digital converters, an acoustic echo canceller (“AEC”), a spatial processor (e.g., one or more multi-channel Wiener filters, one or more other filters, and / or one or more beamformer components), one or more buffers (e.g., one or more circular buffers), one or more wake word engines, one or more speech extractors, and / or one or more other exemplary audio processing components (e.g., components configured to recognize the voices of a particular user or a particular set of users associated with a household). In an example implementation, the audio processing component 220 may include one or more DSPs or one or more modules of a DSP, or may take other forms. In this regard, a particular audio processing component 220 may be configured to have particular parameters (e.g., gain and / or spectral parameters) that can be changed or otherwise adjusted to achieve a particular function. In some implementations, one or more of the audio processing components 220 may be subcomponents of the processor 212.

[0050] In some implementations, the voice processing component 220 can detect and store a user's voice profile associated with the user's account on the MPS 100. For example, the voice profile may be stored and / or compared to a set of command information or variables stored in a data table. The voice profile may include aspects of the tone or frequency of the user's voice and / or other unique aspects of the user's voice, such as those described in the previously-mentioned U.S. Patent Publication No. 2017 / 0242653.

[0051] 2A, the playback device 102 also includes a power component 227. The power component 227 can include at least an external power interface 228, which can be coupled to a power source (not shown), such as via a power cable that physically connects the playback device 102 to an outlet or other external power source. Other power components can include, for example, transformers, converters, and similar components configured to format power.

[0052] In some implementations, the power component 227 of the playback device 102 may further include an internal power source 229 (e.g., one or more batteries) configured to power the playback device 102 without being physically connected to an external power source. With the internal power source 229, the playback device 102 may operate independently of an external power source. In such an implementation, the external power interface 228 may be configured to facilitate charging of the internal power source 229. As previously mentioned, a playback device with an internal power source may be referred to herein as a "portable playback device." Portable playback devices weighing 50 ounces (1,417 g) or less (e.g., between 3 ounces (85.02 g) and 50 ounces (1,417 g), between 5 ounces (141.7 g) and 50 ounces (1,417 g), between 10 ounces (283.4 g) and 50 ounces (1,417 g), between 10 ounces (283.4 g) and 25 ounces (708.5 g), etc.) may be referred to herein as "ultra-portable playback devices." Playback devices that operate using an external power source instead of an internal power source may be referred to herein as "stationary playback devices," although such devices may in fact be mobile within the home or other environments.

[0053] The playback device 102 may further include a user interface 231 that may facilitate user interaction independent of or in conjunction with user interaction facilitated by one or more of the control devices 104. In various embodiments, the user interface 231 includes one or more physical buttons and / or supports a touch-sensitive screen and / or a graphical interface provided on a surface, among other possibilities, for the user to provide direct input. The user interface 231 may further include one or more of lights (e.g., LEDs) and a speaker to provide visual and / or audio feedback to the user.

[0054] 2A, the playback device 102 may also include one or more sensors 209. The sensors 209 may include any suitable sensors, regardless of the value they measure. Examples of suitable sensors include a user-engagement sensor to determine whether a user is wearing or touching the wearable playback device, a microphone or other audio capture device, a camera or other imaging device, an accelerometer, gyroscope, or other motion or activity sensor, a physiological sensor to measure physiological parameters such as heart rate, respiration rate, blood pressure, brain waves, activity level, movement, body temperature, etc. In some implementations, the sensor(s) 209 may be wearable (e.g., the playback device 102 itself is wearable, or the sensor(s) 209 are separate from the playback device 102 while still communicatively coupled to it).

[0055] The playback device 102 may also optionally include a generative media module 211 configured to generate generated media content, either alone or in cooperation with other devices (e.g., other local playback and non-playback devices, remote computing device 154 (FIG. 1C), etc.). As previously described, generated media content may include any media content (e.g., audio, video, audiovisual output, haptic output, or any other media content) that is dynamically created, synthesized, and / or modified by a non-human, rule-based process, such as an algorithm or model, with or without human input. While such a process may be rule-based, it need not be fully deterministic and may instead incorporate randomness or other probabilistic behavior. This creation or modification occurs in real time or near real time for playback. Additionally or alternatively, generated media content may be generated or modified asynchronously (e.g., in advance before playback is requested), and specific items of the generated media content may then be selected for playback at a later time. As used herein, a "generative media module" includes any system capable of generating generative media content based on one or more inputs, whether implemented in software, a physical model, or a combination thereof. In some examples, such generative media content includes novel media content that can be created entirely from scratch or by mixing, assembling, manipulating, or otherwise modifying one or more pieces of existing media content. As used herein, a "generative media content model" includes any algorithm, schema, or set of rules that can be used to generate novel generative media content using one or more inputs (e.g., sensor data, audio captured by an onboard microphone, artist-provided parameters, media segments such as audio clips or samples, etc.).By way of example, a generative media module may generate various generated media content using various generated media content models. In some cases, an artist or other collaborator may interact with, author, and / or update the generated media content model to generate particular generated media content. Although some examples in the description herein refer to audio content, the concepts disclosed herein may, in some examples, be applied to other types of media content, e.g., video, audiovisual, haptic, or other.

[0056] In some examples, the generated media module 211 may utilize one or more input parameters, such as data from the sensor(s) 209, data from other sensors, the capabilities of the playback device (e.g., number and type of transducers, output power, other system architecture), or the location of the device (e.g., location relative to other playback devices, location relative to one or more users). Additional inputs may include device states of one or more devices in the group, such as temperature state (e.g., if a particular device is at risk of overheating, the generated content may be modified to reduce the temperature), battery level (e.g., a portable playback device with a low battery level may have its bass output reduced), bonding state (e.g., whether a particular playback device is configured as part of a stereo pair, bonded with a sub, configured as part of a home theater arrangement, etc.). Other suitable device characteristics or states may likewise be used as inputs for generating the generated media content.

[0057] In the case of a wearable playback device, the position and / or orientation of the playback device relative to the environment can serve as input for the generative media module 211. For example, a spatial soundscape can be provided, where as the user moves through the environment, the corresponding audio changes (e.g., a waterfall in one corner of the room and birds chirping in another). Additionally or alternatively, as the user changes orientation (e.g., looking in a certain direction, tilting their head up or down, etc.), the corresponding audio output changes accordingly.

[0058] User identity (either determined via sensor(s) 209 or otherwise) can also serve as input to the generative media module 211 so that the particular generative media generated by the generative media module 211 is tailored to individual users. For example, when a new user enters a space where generative audio is playing, the user's presence can be detected (e.g., via a proximity sensor, beacon, etc.) and the generated audio can be modified in response. This modification can be based on the number of users (e.g., ambient and meditative audio for one user, relaxing music for two to four users, and party or dance music for four or more users). This modification can also be based on the identity of the user(s) present (e.g., user profiles based on user characteristics, listening history, or other such indicators).

[0059] 2B shows an exemplary housing 230 of a playback device 102 that includes a user interface in the form of a control area 232 on a top 234 of the housing 230. The control area 232 includes buttons 236a-c for controlling audio playback, volume level, and other functions. The control area 232 also includes a button 236d for switching the microphone 222 to either an on or off state.

[0060] 2B, the control area 232 is at least partially surrounded by an opening formed in the top 234 of the housing 230 through which a microphone 222 (not visible in FIG. 2B) receives sounds in the environment of the playback device 102. The microphone 222 can be positioned at various locations along and / or within the top 234 or other areas of the housing 230 to detect sounds from one or more directions relative to the playback device 102.

[0061] As described above, the playback device 102 can be configured as a portable playback device, such as an ultra-portable playback device with an internal power source. FIG. 2C illustrates an example housing 240 of such a portable playback device 102. As shown, the portable playback device housing 240 includes a user interface in the form of a control area 242 on a top 244 of the housing 240. The control area 242 can include a capacitive touch sensor for controlling audio playback, volume level, and other functions. The portable playback device housing 240 can be configured to mate with a dock 246 that connects to an external power source via a cable 248. The dock 246 can be configured to provide power to the portable playback device and recharge its internal battery. In some embodiments, the dock 246 can include one or more sets of conductive contacts (not shown) located on the top of the dock 246 that mate with conductive contacts (not shown) on the bottom of the housing 240. In other embodiments, the dock 246 can provide power to the portable playback device from the cable 248 without using conductive contacts. For example, the dock 246 may wirelessly charge the portable playback device via one or more induction coils integrated into each of the dock 246 and the portable playback device.

[0062] In some embodiments, the playback device 102 can take the form of wired and / or wireless headphones (e.g., over-ear headphones, on-ear headphones, or in-ear headphones). For example, FIG. 2D shows an exemplary housing 250 for such an implementation of the playback device 102. As shown, the housing 250 includes a headband 252 that couples a first earphone 254a to a second earphone 254b. Each of the earpieces 254a and 254b can house any portion of the electronic components within the playback device, such as one or more speakers. Additionally, one or more of the earpieces 254a and 254b can include a control area 258 for controlling audio playback, volume level, and other functions. The control area 258 can be configured with any combination of capacitive touch sensors, buttons, switches, and dials. As shown in FIG. 2D, the housing 250 can further include ear cushions 256a and 256b that are coupled to the earpieces 254a and 254b, respectively. The ear cushions 256a and 256b provide a soft barrier between the user's head and the earpieces 254a and 254b, respectively, to improve user comfort and / or provide acoustic isolation from the surroundings (e.g., passive noise reduction (PNR)). As described above, the playback device 102 can include one or more sensors configured to detect various parameters, optionally including on-ear detection that indicates when the user is wearing or not wearing the headphone device. In some implementations, the wired and / or wireless headphones can be ultra-portable playback devices powered by an internal energy source and weighing less than 50 ounces (1,417 g).

[0063] In some embodiments, the playback device 102 can take the form of in-ear headphones or a hearing aid device. For example, FIG. 2E shows an exemplary housing 260 for such an implementation of the playback device 102. As shown, the housing 260 includes an in-ear portion 262 configured to be positioned in or adjacent to a user's ear and an over-ear portion 264 configured to extend over and behind the user's ear. The housing 260 can house any of the electronic components within the playback device, such as one or more audio transducers, microphones, and audio processing components. Multiple control areas 266 can facilitate user input for controlling audio playback, volume levels, noise cancellation, pairing with other devices, and other functions. The control area 258 can consist of any combination of one or more buttons, switches, dials, capacitive touch sensors, and the like. As described above, the playback device 102 can include one or more sensors configured to detect various parameters, optionally including in-ear detection that indicates when a user is wearing or not wearing the in-ear headphone device.

[0064] It should be understood that the playback device 102 can take the form of other wearable devices aside from headphones. Wearable devices include devices configured to be worn on a part of the subject (e.g., head, neck, torso, arm, wrist, finger, leg, ankle, etc.). For example, the playback device 102 can take the form of eyeglasses including a frame front (e.g., configured to hold one or more lenses), a first temple rotatably coupled to the frame front, and a second temple rotatably coupled to the frame front. In this example, the eyeglasses can include one or more transducers integrated into at least one of the first and second temples and configured to project sound toward the subject's ears.

[0065] While specific implementations of playback devices and network microphone devices are described above with reference to Figures 2A-2E, numerous device configurations exist, including, but not limited to, those without a UI, microphones in different locations, multiple microphone arrays arranged in different configurations, and / or other configurations appropriate to the requirements of a given application. For example, the UI and / or microphone array may be implemented in other playback devices and / or computing devices other than those described herein. Furthermore, while specific examples of playback device 102 are described with reference to MPS 100, those skilled in the art will recognize that the playback devices described herein can be used in a variety of different environments, including, but not limited to, environments with more and / or fewer elements, without departing from the invention. Similarly, the MPS described herein can be used with a variety of different playback devices.

[0066] By way of example, SONOS, Inc. currently offers (or has offered) for sale certain playback devices that may implement some of the embodiments disclosed herein, including "SONOS ONE," "FIVE," "PLAYBAR," "AMP," "CONNECT:AMP," "PLAYBASE," "BEAM," "ARC," "CONNECT," "MOVE," "ROAM," and "SUB." Any other past, present, and / or future playback devices may additionally or alternatively be used to implement the playback devices of the exemplary embodiments disclosed herein. Furthermore, it should be understood that playback devices are not limited to the examples shown in FIGS. 2A-2D or to SONOS products. For example, playback devices may be integrated with other devices or components, such as televisions, lighting fixtures, or other indoor or outdoor devices.

[0067] b. Playback device configuration example 3A-3E show example playback device configurations. Referring first to FIG. 3A, in some examples, a single playback device may belong to a zone. For example, playback device 102c (FIG. 1A) on the patio may belong to Zone A. In some embodiments described below, multiple playback devices may be "bonded" to form a "bonded pair" that combines to form a single zone. For example, playback device 102f (FIG. 1A) named "Bed 1" in FIG. 3A may be bonded with playback device 102g (FIG. 1A) named "Bed 2" in FIG. 3A to form Zone B. The bonded playback devices may have different playback responsibilities (e.g., channel responsibilities). In other embodiments described below, multiple playback devices may be merged to form a single zone. For example, playback device 102d named "Bookshelf" may be bonded with playback device 102m named "Living Room" to form Zone C. The bonded playback devices 102d and 102m may not be assigned different playback responsibilities. That is, the integrated playback devices 102d and 102m can each play audio content in the same way as if they were not integrated, other than playing audio content in synchronization.

[0068] For control purposes, each zone of the MPS 100 can be represented as a single user interface (“UI”) entity. For example, as displayed by the control device 104, zone A is provided as a single entity called “portable,” zone B is provided as a single entity called “stereo,” and zone C is provided as a single entity called “living room.”

[0069] In various embodiments, a zone can be named after one of the playback devices that belong to that zone. For example, Zone C can be named after Living Room Device 102m (as shown). In another example, Zone C can instead be named after Bookshelf Device 102d. In yet another example, Zone C could be named after a combination of Bookshelf Device 102d and Living Room Device 102m. The selected name can be selected by the user via input on Control Device 104. In some embodiments, a zone is given a name that is different from the devices that belong to that zone. For example, Zone B in FIG. 3A is named “Stereo,” but none of the devices in Zone B have this name. In one aspect, Zone B is a single UI entity representing a single device named “Stereo” that is composed of “Bed 1” and “Bed 2.” In one embodiment, Bed 1 device may be playback device 102f in Master Bedroom 101h (FIG. 1A), and Bed 2 device may be playback device 102g, also in Master Bedroom 101h (FIG. 1A).

[0070] As mentioned above, coupled playback devices may have different playback responsibilities, such as playback responsibilities for specific audio channels. For example, as shown in FIG. 3B, Bed 1 and Bed 2 devices 102f and 102g may be coupled to create or enhance a stereo effect for audio content. In this example, Bed 1 playback device 102f may be configured to play the left channel audio component, and Bed 2 playback device 102g may be configured to play the right channel audio component. In some implementations, such stereo coupling may be referred to as "pairing."

[0071] Furthermore, playback devices configured to be coupled can have additional and / or different speaker drivers. As shown in FIG. 3C, a playback device 102b labeled "Front" can be coupled to a playback device 102k labeled "SUB." The Front device 102b reproduces mid- to high-frequency sounds, while the SUB device 102k reproduces low-frequency sounds, e.g., as a subwoofer. When uncoupled, the Front device 102b can be configured to render a full range of frequencies. As another example, FIG. 3D shows the Front device 102b and the SUB device 102k further coupled to the right playback device 102a and the left playback device 102j, respectively. In some implementations, the right device 102a and the left device 102j can form surround or "satellite" channels in a home theater system. The combined playback devices 102a, 102b, 102j, and 102k can form a single zone D (FIG. 3A).

[0072] In some implementations, playback devices may also be “merged.” In contrast to a combined playback device, a merged playback device is not assigned playback responsibilities and can each render the full range of audio content that each playback device is capable of. However, merged devices may be represented as a single UI entity (i.e., a zone, as described above). For example, FIG. 3E shows playback devices 102d and 102m in the living room merged, such that they are represented by a single UI entity for Zone C. In one embodiment, playback devices 102d and 102m may play audio synchronously, while outputting the full range of audio content that each playback device 102d and 102m is capable of rendering.

[0073] In some embodiments, a standalone NMD may exist in a zone by itself. For example, NMD 103h in FIG. 1A is named "Closet" and forms Zone I in FIG. 3A. NMDs may also be combined or integrated with other devices to form zones. For example, NMD device 103f named "Island" is combined with playback device 102i "Kitchen" to form Zone F named "Kitchen." Additional details regarding assigning NMDs and playback devices as designated or default devices are described, for example, in the previously discussed U.S. Patent Publication No. 2017-0242653. In some embodiments, a standalone NMD is not assigned to a zone.

[0074] Individual, combined, and / or merged zones of devices may be arranged to form a set of playback devices that play audio in synchronization. Such a set of playback devices may be referred to as a “group,” “zone group,” “sync group,” or “playback group.” In response to input provided via control device 104, playback devices may be dynamically grouped and ungrouped to form new or different groups that play audio content in synchronization. For example, with reference to FIG. 3A , zone A may be grouped with zone B to form a zone group containing the playback devices of the two zones. As another example, zone A may be grouped with one or more other zones C-I. Zones A-I may be grouped and ungrouped in many ways. For example, three, four, five, or more (e.g., all) of zones A-I may be grouped. When grouped, individual and / or combined zones of playback devices may play audio in synchronization with one another, as described in the previously discussed U.S. Patent No. 8,234,395. Grouped devices and combined devices are exemplary types of associations between portable and stationary playback devices that may be created in response to a trigger event, as described above and in further detail below.

[0075] In various implementations, zones within an environment can be assigned specific names, which can be default names for zones within a zone group or combinations of the names of zones within a zone group, such as "Dining Room + Kitchen," as shown in Figure 3A. In some embodiments, zone groups are given unique names selected by the user, such as "Nick's Room," as shown in Figure 3A. The name "Nick's Room" might be a name selected by the user over a previous name for the zone group, such as Master Bedroom.

[0076] 2A, certain data may be periodically updated and stored in memory 213 as one or more state variables used to describe the state of a playback zone, playback device, and / or its associated zone group. Memory 213 may also contain data related to the state of other devices in media playback system 100, which may be shared between devices from time to time so that one or more devices have the most current data related to the system.

[0077] In some embodiments, the memory 213 of the playback device 102 can store instances of various state-related variable types. The variable instances may be stored with an identifier (e.g., a tag) corresponding to the type. For example, one identifier may be a first type "a1" to identify a playback device for a zone, a second type "b1" to identify playback devices that may be coupled to the zone, and a third type "c1" to identify a zone group to which the zone belongs. As a related example, in FIG. 1A , an identifier associated with the patio may indicate that the patio is the only playback device for a particular zone and is not included in any zone group. An identifier associated with the living room may indicate that the living room is not grouped with other zones and includes coupled playback devices 102a, 102b, 102j, and 102k. An identifier associated with the dining room may indicate that the dining room is part of the dining room + kitchen group and that devices 103f and 102i are coupled. An identifier associated with the kitchen may indicate the same or similar information because the kitchen is part of the "dining room + kitchen" zone group. Other examples of zone variables and identifiers are described below.

[0078] In yet another example, the MPS 100 may include variables or identifiers representing other associations between zones and zone groups, such as an identifier associated with an "area," as shown in FIG. 3A. An area may include a zone group and / or a cluster of zones not included in a zone group. For example, FIG. 3A illustrates a first area labeled "Area 1" and a second area labeled "Area 2." The first area includes zones or zone groups for the patio, den, dining room, kitchen, and bathroom. The second area includes zones or zone groups for the bathroom, Nick's Room, bedroom, and living room. In one respect, "area" may be used to refer to zone groups and / or clusters of zones that share one or more zones and / or zone groups with other clusters. In this respect, such areas are distinct from zone groups that do not share any zones with other zone groups. Further examples of techniques for implementing areas are described, for example, in U.S. Patent Publication No. 2018-0107446, published April 19, 2018, and entitled "Room Association Based on Name," and U.S. Patent No. 8,483,853, filed September 11, 2007, and entitled "Controlling and manipulating groupings in a multi-zone media system," each of which is incorporated herein by reference in its entirety. In some embodiments, MPS 100 may not implement areas, in which case the system may not store variables associated with areas.

[0079] The memory 213 may be further configured to store other data. Such data may relate to audio sources accessible by the playback device 102 or playback queues with which the playback device (or other playback device(s)) may be associated. In embodiments described below, the memory 213 is configured to store a set of command data for selecting a particular VAS when processing audio input.

[0080] In operation, one or more playback zones in the environment of FIG. 1A may each be playing different audio content. For example, a user may be grilling in the patio zone and listening to hip hop music played by playback device 102c, while another user may be preparing food in the kitchen zone and listening to classical music played by playback device 102i. In another example, a playback zone may synchronize with another playback zone to play the same audio content. For example, a user may be in the office zone and playback device 102n may be playing the same hip hop music that playback device 102c is playing in the patio zone. In such a case, playback devices 102c and 102n may play hip hop in synchronization so that a user can move between different playback zones and enjoy the audio content being played at high volume seamlessly (or at least substantially seamlessly). Synchronization between playback zones may be achieved in a manner similar to synchronization between playback devices, as described in the previously mentioned U.S. Patent No. 8,234,395.

[0081] As described above, the zone configuration of the MPS 100 can be dynamically changed. In this manner, the MPS 100 can support multiple configurations. For example, if a user physically moves one or more playback devices into or out of a zone, the MPS 100 may be reconfigured to accommodate the change. For example, if a user physically moves playback device 102c from the patio zone to the office zone, the office zone would include both playback devices 102c and 102n. In some cases, the user may pair or group the moved playback device 102c with the office zone and / or rename the player in the office zone, for example, using either the control device 104 and / or voice input. As another example, if one or more playback devices 102 are moved to a particular space in a home environment that is not already a playback zone, the moved playback device(s) may be renamed or associated with the playback zone of the particular space.

[0082] Additionally, different playback zones of the MPS 100 can be dynamically combined into zone groups or split into individual playback zones. For example, a “dining room” zone and a “kitchen” zone can be combined into a zone group for a dinner party, with playback devices 102i and 102l synchronously rendering audio content. As another example, the combined playback devices in a den zone can be split into (a) a television zone and (b) a separate listening zone. The television zone can include the front playback device 102b. The listening zone can include right, left, and sub playback devices 102a, 102j, and 102k, which can be grouped, paired, or combined as described above. Splitting the den zone in this manner allows one user to listen to music in a listening zone in one area of ​​the living room space, while another user watches television in another area of ​​the living room space. In a related example, users can control the den zone using either NMD 103a or 103b (FIG. 1B) before it is separated into a television zone and a listening zone. Once separated, the listening zone is controlled by a user in proximity to, for example, NMD 103a, and the television zone is controlled by a user in proximity to, for example, NMD 103b. However, as noted above, any of the NMDs 103 may be configured to control various playback and other devices of the MPS 100.

[0083] c. Control Device Examples FIG. 4A is a functional block diagram illustrating selected aspects of the control device 104 of the MPS 100 of FIG. 1A. A control device according to some embodiments of the present invention can be used in a variety of systems, such as (but not limited to) an MPS as described in FIG. 1A. Such a control device may also be referred to herein as a “control device” or a “controller.” The control device illustrated in FIG. 4A may include components generally similar to certain components of the network devices described above, such as a processor 412, memory 413 storing program software 414, at least one network interface 424, and one or more microphones 422. As one example, the control device may be a dedicated controller for the MPS 100. As another example, the control device may be a network device on which media playback system control application software is installed, such as an iPhone®, iPad®, other smartphone, tablet, or network device (e.g., a network computer such as a PC or Mac™).

[0084] The memory 413 of the control device 104 may be configured to store control application software and other data related to the MPS 100 and / or users of the system 100. The memory 413 may be loaded with software 414 instructions executable by the processor 412 to implement certain functions, such as facilitating user access, control, and / or configuration of the MPS 100. The control device 104 may be configured to communicate with other network devices via a network interface 424, which may take the form of a wireless interface, as described above.

[0085] In one example, system information (e.g., state variables, etc.) may be communicated between the control device 104 and other devices via the network interface 424. For example, the control device 104 may receive settings for playback zones and zone groups within the MPS 100 from a playback device, an NMD, or another network device. Similarly, the control device 104 may transmit such system information to the playback device or another network device via the network interface 424. In some cases, the other network device may be another control device.

[0086] The control device 104 may also communicate playback device control commands, such as volume control and audio playback control, to the playback devices via the network interface 424. As alluded to above, configuration changes to the MPS 100 may also be performed by a user using the control device 104. Configuration changes may include adding / removing one or more playback devices to / from a zone, adding / removing one or more zones to / from a zone group, forming a combined or integrated player, separating one or more playback devices from a combined or integrated player, etc.

[0087] As shown in FIG. 4A , the control device 104 may also generally include a user interface 440 configured to facilitate user access and control of the MPS 100. The user interface 440 may include a touchscreen display or other physical interface configured to provide various graphical controller interfaces, such as the controller interfaces 440a and 440b shown in FIGS. 4B and 4C . Referring simultaneously to FIGS. 4B and 4C , the controller interfaces 440a and 440b include a playback control area 442, a playback zone area 443, a playback status area 444, a playback queue area 446, and a source area 448. The illustrated user interface is merely one example of an interface that may be provided on a network device, such as the control device shown in FIG. 4A , and that a user may access to control a media playback system, such as the MPS 100. Other user interfaces of various formats, styles, and interaction sequences may also be implemented on one or more network devices to provide comparable control access to a media playback system.

[0088] Playback control area 442 (FIG. 4B) may include selectable icons (e.g., by touch or using a cursor) that, when selected, cause playback devices in the selected playback zone or zone group to play or pause, fast forward, rewind, skip next, skip previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit crossfade mode, etc. Playback control area 442 may also include selectable icons that, when selected, change equalization settings and / or playback volume, among other possibilities.

[0089] Play zones area 443 (FIG. 4C) may include representations of play zones within MPS 100. Play zones area 443 may also include representations of zone groups, such as the "dining room + kitchen" zone group, as shown. In some embodiments, the graphical representation of play zones may be selectable to invoke additional selectable icons for managing or configuring play zones within MPS 100, such as creating combined zones, creating zone groups, separating zone groups, renaming zone groups, among other possibilities.

[0090] For example, as shown, a "group" icon may be provided within each graphical representation of a playback zone. The "group" icon within the graphical representation of a particular zone is selectable to display options for selecting and grouping one or more other zones within MPS 100 with the particular zone. Once grouped, the playback devices of the zones grouped with the particular zone are configured to play audio content in synchronization with the playback devices of the particular zone. Similarly, a "group" icon may be provided within the graphical representation of a zone group. In this case, selecting the "group" icon may display options for removing one or more zones within the zone group from the zone group. Other interactions and implementations for grouping and ungrouping zones via the user interface are also possible. The representation of the playback zones in playback zones area 443 (FIG. 4C) dynamically updates as the configuration of the playback zones or zone groups changes.

[0091] Playback status area 444 (FIG. 4B) may include a graphical representation of audio content currently playing, previously played, or next scheduled to play in a selected playback zone or zone group. The selected playback zone or zone group may be visually distinguishable on the controller interface, such as in playback zone area 443 and / or playback status area 444. The graphical representation may include track title, artist name, album name, album year, track length, and / or other relevant information that is useful for a user to know when controlling MPS 100 via the controller interface.

[0092] The play queue area 446 may include a graphical representation of audio content in a play queue associated with a selected playback zone or zone group. In some embodiments, each playback zone or zone group may be associated with a play queue that includes information corresponding to zero or more audio items for playback by the playback zone or zone group. For example, each audio item in the play queue may include a Uniform Resource Identifier (URI), Uniform Resource Locator (URL), or other identifier that a playback device in the playback zone or zone group can use to search for and / or obtain the audio item from a local or network audio content source, which can then be played by the playback device.

[0093] In one example, a playlist may be added to the play queue, where information corresponding to each audio item in the playlist may be added to the play queue. In another example, the audio items in the play queue may be saved as a playlist. In a further example, the play queue may be empty, or it may exist but be "unused" if a playback zone or zone group is playing continuous streaming audio content that continues to play until stopped, such as Internet radio, rather than individual audio items with play times. In alternative embodiments, the play queue may contain Internet radio and / or other streaming audio content items and be "in use" when a playback zone or zone group is playing those items. Other examples are possible.

[0094] When a playback zone or zone group is "grouped" or "ungrouped," the playback queues associated with the affected playback zones or zone groups are cleared or reassociated. For example, if a first playback zone containing a first playback queue is grouped with a second playback zone containing a second playback queue, the established zone group can have an associated playback queue that is initially empty, contains audio items from the first playback queue (e.g., if the second playback zone is added to the first playback zone), contains audio items from the second playback queue (e.g., if the first playback zone is added to the second playback zone), or contains a combination of audio items from both the first and second playback queues. If the established zone group is subsequently ungrouped, the resulting first playback zone may be reassociated with the previous first playback queue or associated with a new playback queue, in the latter case either empty or containing audio items from the playback queues associated with the established zone group before the established zone group was ungrouped. Similarly, the resulting second playback zone may be reassociated with the previous second playback queue, or may be associated with an empty new playback queue, or may contain audio items from the playback queue that was associated with the established zone group before the established zone group was ungrouped. Other examples are possible.

[0095] 4B and 4C, the graphical representation of audio content in the play queue area 446 (FIG. 4B) may include track title, artist name, track length, and / or other relevant information related to the audio content in the play queue. In one example, the graphical representation of audio content may be selectable to display additional selectable icons for managing and / or manipulating the play queue and / or the audio content represented in the play queue. For example, representative audio content may be removed from the play queue, moved to a different position in the play queue, selected for immediate play, or played after currently playing audio content. A play queue associated with a playback zone or zone group may be stored in memory on one or more playback devices included in the playback zone or zone group, on playback devices not included in the playback zone or zone group, and / or on other designated devices. Playback of such a play queue involves one or more playback devices playing the media items in the queue, possibly in sequential or random order.

[0096] The source area 448 may include a graphical representation of selectable audio content sources and / or selectable voice assistants associated with the corresponding VAS. VASs may be selectively assigned. In some examples, multiple VASs, such as Amazon's Alexa and Microsoft's Cortana, may be activated by the same NMD. In some embodiments, a user may exclusively assign a VAS to one or more NMDs. For example, a user may assign a first VAS to one or both of the living room NMDs 102a and 102b shown in FIG. 1A and a second VAS to the kitchen NMD 103f. Other examples are possible.

[0097] d. Audio Content Source Examples The audio sources in source area 448 are audio content sources from which audio content is obtained and played by a selected playback zone or zone group. One or more playback devices within a zone or zone group may be configured to obtain audio content for playback from various available audio content sources (e.g., according to a URI or URL corresponding to the audio content). As one example, audio content may be obtained directly from an audio content source supported by the playback device (e.g., via a line-in connection). In another example, audio content may be provided to the playback device over a network via one or more other playback or network devices. As described in more detail below, in some embodiments, audio content may be provided by one or more media content services.

[0098] Examples of audio content sources may include memory on one or more playback devices within a media playback system such as MPS 100 of FIGS. 1A-1C, a local music library on one or more networked devices (e.g., a control device, a network-enabled personal computer, or a network-attached storage (“NAS”)), a streaming audio service that provides audio content over the Internet (e.g., a cloud-based music service), or an audio source connected to the media playback system via a line-in connection on a playback device or network device.

[0099] In some embodiments, audio content sources can be added or removed from a media playback system such as MPS 100. In one example, audio item indexing may be performed whenever one or more audio content sources are added, removed, or updated. Audio item indexing may include scanning for identifiable audio items in all folders / directories shared on a network accessible by playback devices in the media playback system and, for each identifiable audio item found, generating or updating an audio content database consisting of metadata (e.g., title, artist, album, track length, etc.) and other associated information such as a URI or URL. Other examples for managing and maintaining audio content sources are also possible.

[0100] III. Example of generated audio playback using a wearable playback device Generative audio content enables the creation and delivery of audio content tailored to a particular user, a particular environment, and / or a particular time. In the case of a wearable playback device, the generative audio content can be further tailored based on parameters detected via the wearable playback device, and playback can be controlled and distributed among various playback devices in an environment based on contextual data. For example, a wearable playback device can detect that it is being worn by a user and automatically begin playing generative audio content, including a generative musical composition based at least in part on one or more media content stems and / or audio cues derived from the contextual data. Contextual data can include information about the user's environment (e.g., time of day, temperature, circadian rhythm, humidity, number of nearby people, ambient light level), or other types of indicators (e.g., doorbell, alarm, event). Such contextual data can be obtained via on-board sensors on the wearable playback device, sensors associated with other playback devices in the environment, or any other suitable sensor data source. In certain examples, the wearable playback device can identify a particular user when the playback device is placed on the user's head, and the wearable playback device can further tailor the generated media content to the user's profile, current or desired emotional state, and / or other biometric data. Furthermore, in some cases, playback of the generated audio content is dynamically swapped or switched between playback solely through the wearable playback device and alternative or simultaneous playback through one or more out-loud playback devices in the listening environment. In some examples, the wearable playback device is played through one or more out-loud playback devices, and optionally provides one or more contextual inputs to the generated soundscape played through the wearable playback device. Consider, for example, a scenario in which one or more listeners desire to monitor biometric data (e.g., heart rate, respiration rate, body temperature, blood glucose level, blood oxygen saturation) of the wearer of the wearable playback device.Because the wearer may be ill, elderly, a child, or for other reasons have abilities that differ from a typical out-loud listener, one or more generative soundscapes can be generated and played using one or more biometric parameters of the wearer.

[0101] a. Creation and playback of generative audio content using wearable playback devices 5A-5D illustrate an exemplary method for generative audio playback via a wearable audio playback device. Referring to FIG. 5A, method 500 begins at block 502 with on-ear detection. For example, on-board sensors can determine when the wearable playback device is positioned on or relative to the user's head (e.g., with ear cups over the user's ears or earbuds in the user's ears). The on-board sensors can take any suitable form, such as optical proximity sensors, capacitive or inductive touch sensors, gyroscopes, accelerometers, or other motion sensors.

[0102] Upon on-ear detection, method 500 proceeds to block 504 and automatically initiates playback of the generated audio content via the wearable playback device. As previously described, generated media content (e.g., generated audio content) can be generated via an on-board generated media module present on the wearable playback device, via generated media module(s) present on another local playback or computing device (e.g., accessible via a local area network (e.g., WiFi) or a direct wireless connection (e.g., Bluetooth)). Additionally or alternatively, the generated media content may be generated by one or more remote computing devices, optionally using one or more input parameters provided by the wearable playback device and / or a media playback system including the wearable playback device. In some examples, the generated media content may be generated by any combination of these devices.

[0103] Once the on-ear condition is no longer detected at block 506 (e.g., via one or more on-board sensors), playback of the generated audio content stops at block 508. Optionally, playback of the generated media content may automatically transition or swap to playback via one or more other playback devices in the environment, such as one or more out-loud playback devices. Playback control, including such swapping, is described in more detail below with respect to Figures 10A-10D.

[0104] 5A, the wearable playback device is configured to automatically begin playing the generated audio content when worn by a user and to automatically end playing the generated audio content when removed by the user. Optionally, this process can be performed without tethering or otherwise connecting the wearable playback device to another device.

[0105] In various examples, the generated audio content can be a soundscape tailored to the user and / or the user's environment. For example, the generated audio can be responsive to the user's current emotional state (e.g., determined via one or more sensors in the wearable playback device) and / or contextual data (e.g., information about the user's environment or home, external data such as events outside the user's environment, the time of day, etc.).

[0106] 5B illustrates another example method 510 that begins at block 512 with playing generated audio via a wearable playback device. At block 514, a Bluetooth connection to an audio source (e.g., a mobile phone, a tablet, etc.) or a WiFi audio source is detected. After this detection, method 510 transitions from generated audio playback to source audio playback at block 516. Optionally, this transition can occur only when source audio playback begins, thereby preventing the user from being left with unnecessary silence.

[0107] In some examples, the transition may include a crossfade of a predetermined duration (e.g., 5 seconds, 10 seconds, 20 seconds, 30 seconds) to ensure a gradual transition between the generated audio content and the source audio. The transition may also be based on other factors, such as transients or "drops" in the content, or may be staggered by frequency bands and the energy detected within those bands. In some cases, the source audio type (e.g., music, speech, telephone) may also be detected, and the crossfade time adjusted accordingly. For example, the crossfade may be long (e.g., 10 seconds) for music, medium (e.g., 5 seconds) for speech, and short (e.g., 0, 1, 2 seconds) for telephone. Optionally, the transition may include starting active noise control (ANC) if it was not already in use during playback, or ending active noise control (ANC) if it was previously in use.

[0108] 5C shows an example method 520. At block 522, the Bluetooth, WiFi, or other wirelessly connected audio source is no longer detected at the wearable playback device. At block 524, the wearable playback device automatically transitions to playing the generated audio content. This transition may similarly include toggling the crossfade and / or ANC features described above.

[0109] Optionally, as shown in block 526, previously played audio from a Bluetooth or WiFi source (e.g., block 516 of FIG. 5B ) can be used as an input parameter (e.g., stem or seed) for the generative media module so that when playback of the generated audio content resumes in block 524, the generated audio content includes certain characteristics or is otherwise based at least in part on the previous source audio, thereby enhancing the sense of continuity between different media content being played.

[0110] 5D illustrates an example method 530 and accompanying scenario for playing generated audio content through multiple playback devices. Method 530 begins at block 532 with initiating a group generated audio playback session. For example, a wearable playback device may be grouped for synchronized playback with one or more additional playback devices (which may include one or more out-of-loud playback devices). In this configuration, the various playback devices can synchronize and play the generated media content. In the example illustrated in FIG. 5D, wearable device 540 is grouped with a second wearable playback device 542 and out-of-loud playback devices 544, 546, and 548.

[0111] At block 534, a first playback device generates the generated audio content. In various examples, the first playback device may be a wearable playback device (e.g., playback device 540), an out-loud playback device (e.g., playback device 546), or a combination of playback devices that cooperate to generate the generated audio content. As noted elsewhere, the generated audio content may be based at least in part on sensor data and / or contextual data, which may be derived at least in part from the playback device itself.

[0112] The method 530 continues by transmitting the generated audio content to additional playback devices in the environment at block 536. The various playback devices can then play the generated audio content in synchronization with one another.

[0113] In some examples, wearable playback device 542 may be selected as the group coordinator for a synchronized group, especially if wearable playback device 542 is already playing generated audio content. However, in other examples, playback device 548 may be selected as the group coordinator because, for example, it has the most computing power, it is connected to a power source, it has better network hardware or network connections, etc.

[0114] In some examples, an out-loud playback device 548 (e.g., a subwoofer) may be coupled to the out-loud playback device 546 (e.g., a sound bar), and low-frequency audio content may be played through the out-loud playback device 548 rather than the audio being played through the out-loud playback device 546. For example, a listener of the wearable playback device 540 and / or the wearable playback device 542 may desire to take advantage of the low-frequency capabilities provided by a subwoofer, such as the out-loud playback device 548, while listening to the generated audio content through the wearable playback devices 540, 542. In some implementations, the sub can play low-frequency content intended to be listened to along with multiple, but different, items of high-frequency content (e.g., a single sub and sub-content channel can provide low-frequency content for both headphone and out-loud listeners, even if the high-frequency content of each is different).

[0115] 6 is a schematic diagram of an exemplary distributed generative media playback system 600. As shown, an artist 602 can provide multiple media segments 604 and one or more generative content models 606 to a stored generative media module 214 via one or more remote computing devices. A media segment can correspond, for example, to a particular audio segment or seed (e.g., individual notes or chords, a short n-bar track, non-musical content, etc.). In some examples, the generative content model 606 can also be provided by the artist 602. This can include providing the entire model, or the artist 602 can provide input to the model 606 by, for example, modifying or adjusting certain aspects (e.g., tempo, melodic constraints, harmonic complexity parameters, chord change density parameters, etc.). Furthermore, this process can include seamless loops that take segments of long stems and recombine them infinitely.

[0116] The generative media module 211 may receive both a media segment 604 and one or more input parameters 603 (as described elsewhere herein). Based on these inputs, the generative media module 211 may output generated media. As shown in FIG. 6, the artist 602 may optionally audition the generative media module 211, for example, by receiving exemplary outputs based on inputs (e.g., media segments 604 and / or generative content models 606) provided by the artist 602. In some cases, the audition may play variations of the generated media content to the artist 602 in response to a variety of different input parameters (e.g., one version corresponding to a high energy level intended to create an exciting or uplifting effect, another version corresponding to a low energy level intended to create a calming effect, etc.). Based on the output from this auditioning step, the artist 602 may dynamically update settings of the media segment 604 and / or generative content model 606 until a desired output is achieved.

[0117] In the illustrated example, there may be an iteration in block 608 every n hours (or minutes, days, etc.) during which the generated media module 214 may generate multiple different versions of the generated media content. In the illustrated example, there are three versions: version A in block 610, version B in block 612, and version C in block 614. These outputs are stored (e.g., via a remote computing device) as generated media content 616. A particular one of the versions (version C in this example as block 618) may be transmitted (e.g., streamed) to the local playback device 102a for playback.

[0118] Although three versions are shown here by way of example, in practice there may be many more versions of generated media content generated via a remote computing device. The versions may vary along several different dimensions, such as being suitable for different energy levels, for different intended tasks or activities (e.g., studying versus dancing), for different times of day, or any other suitable variation.

[0119] In the illustrated example, the playback device 102a can periodically request a particular version of the generated media content from a remote computing device. Such a request can be based, for example, on user input (e.g., user selection via a controller device), sensor data (e.g., the number of people present in a room, background noise level, etc.), or other suitable input parameters. As illustrated, the input parameters 603 can optionally be provided to (or detected by) the playback device 102a. Additionally or alternatively, the input parameters 603 can be provided to (or detected by) the remote computing device 106. In some examples, the playback device 102a transmits the input parameters to the remote computing device 106, which then provides the appropriate version to the playback device 102a without the playback device 102a specifically requesting a particular version.

[0120] Outloud playback device 102a can transmit generated media content 616 to wearable playback device 102o via WiFi, Bluetooth, or other suitable wireless connection. Optionally, wearable playback device 102o can also provide one or more input parameters (e.g., sensor data from wearable playback device 102o) that can be used as input for generated media module 211. Additionally or alternatively, wearable playback device 102o can receive one or more input parameters 603 that can be used to modify the playback of the generated audio content via wearable playback device 102o.

[0121] 7 is a schematic diagram of an example system 700 for generating and playing back generated media content. System 700 may be similar to system 600 described above with respect to FIG. 6, except that generated media content 616 (e.g., version C 618) may be transmitted directly from remote computing device(s) to wearable playback device 102o. As previously mentioned, the generated media content may be further distributed from wearable playback device 102o to additional playback devices in the environment, including other wearable playback devices and / or out-of-home playback devices.

[0122] b. Positioning and customization for generated media playback In various examples, the generation, delivery, and playback of generated media content can be controlled and / or adjusted based on location information, user identity, or other such parameters. FIGS. 8A-8D illustrate example methods for playing generated audio based on location. With reference to FIG. 8A, method 800 begins with on-ear detection at block 802. Optionally, after on-ear detection, method 800 proceeds to block 804 to identify the user. User identification can be performed using any suitable technology, such as voice recognition, a fingerprint touch sensor, entry of a user-specific code via user input, facial recognition via an imaging device, or any other suitable technology. In some cases, user identity can be inferred based on connected Bluetooth device(s) (e.g., if the wearable playback device is connected to John Smith's mobile phone, John Smith is identified as the user after on-ear detection). Optionally, the user identity can be a pseudo-identity, in which a separate profile of preferences (desired values) is stored but does not correspond to a single user account or real identity.

[0123] At block 806, the wearable playback device plays back the generated audio content specific to the user's location and, optionally, the identified user. The location can be determined using any suitable localization technique, such as exchanging localization signals with other devices in the environment (e.g., UWB localization, WiFi RSSI evaluation, acoustic localization signals, etc.). In some examples, the wearable playback device can include a GPS component configured to provide absolute location data. As shown, process 800 can branch to methods 820, 830, and / or 840, described below with respect to FIGS. 8B, 8C, and 8D. Other examples of location determination are described in the accompanying appendices.

[0124] In some examples, the generated audio content generated and played in block 806 is based on certain real-time contextual information (e.g., temperature, light, humidity, time of day, the user's emotional state) in addition to location information. If the user's identity has been determined, the generated audio content may be generated according to one or more user preferences or other input parameters specific to that user (e.g., the user's listening history, favorite artists, etc.). In some cases, the wearable playback device may begin playing the generated audio content only if the identified user is authorized to use or associated with the wearable device.

[0125] If, at decision block 808, the location has not changed, the method returns to block 806 to continue playing the generated audio content based on the location, i.e., position, and optionally, the user's identification. If, at decision block 808, the location has changed (e.g., as determined via sensor data), method 800 continues to block 810 by identifying the new location and, at block 812, playing the specific generated audio content based on the new location. For example, within a home environment, moving from the kitchen to a home office might automatically transition the generated audio content to more peaceful, focus-enhancing audio content, while moving to a home gym might transition the generated audio content to more upbeat, high-energy audio content. The transition method may have a significant impact on the user experience. In most cases, a slow / smooth transition is preferred, but transitions may be more complex (e.g., increasing the sense of spaciousness so that the sound feels contained within a given room, or partially blending with the audio of nearby users when in close proximity).

[0126] In certain examples, the generated audio content may be generated based, at least in part, on a theme, mood, and / or pleasant scenery associated with a particular location. For example, a room may be themed (e.g., a forest room, a beach room, a tropical jungle room, a workout room), and the generated audio content generated and played is based on the soundscape associated with that room. The generated audio content played via the wearable playback device may be identical to, substantially similar to, or otherwise related to audio content normally played by one or more out-loud devices in the room. In some examples, context-based audio cues may be layered over the generated audio content associated with the room.

[0127] 8B shows an example method 820 in which audio and / or sounds within the current location are detected at block 822. For example, the wearable playback device may include one or more microphones (and / or other microphones in the environment may capture sound data and transmit it to the wearable playback device). At block 824, the detected audio and / or sounds may be used as input for generated audio content. Outloud audio played via a nearby playback device, ambient noises and sounds in the user's environment, and / or audio played nearby another user via another wearable playback device may be used as input for a generative media module that generates generated audio content that is responsive to, matches, and / or is at least partially influenced by the detected audio in the environment.

[0128] 8C shows an example method 830 for controlling light sources in a manner responsive to generated audio content. Method 830 begins at block 832 with determining available light sources based on the identified location. (These may be, for example, remotely controlled internet-connected light bulbs, televisions, or other displays.) The light sources may be distributed in particular rooms or other locations within the environment. At block 834, current generated audio content characteristics are mapped to available light source(s). Then, at block 836, the light source(s) adjust their respective lighting parameter(s) based on the mapped characteristic(s).

[0129] Characteristics of the generated audio content that can be mapped to available light sources include, for example, mood, tempo, individual notes, chords, etc. Based on the mapped generated audio content characteristics, individual light sources can adjust their corresponding parameters (e.g., color, brightness / intensity, color temperature, etc.) accordingly.

[0130] FIG. 8D shows an example method 840 for detecting location and determining room acoustics of a particular listening environment (e.g., an environment in which a listener of a wearable playback device is located). The method 840 begins with detecting a location identifier at block 842, and at block 844, the method 840 includes determining location based on the detected identifier. This may include determining where the user is currently located using an appropriate sensor modality or combination of modalities (e.g., UWB localization, acoustic localization, BLE localization, ultrasonic localization, motion sensor data (e.g., IMU), wireless RSSI localization, etc.). In some examples, the user's location is mapped to a predefined room or space and stored via the media playback system for purposes of managing synchronized audio playback throughout the user's environment.

[0131] If, at decision block 846, no acoustic devices in the room are determined, the process may proceed to method 800 shown in FIG. 8A. If acoustic devices in the room are determined, method 840 proceeds to decision block 848. If devices in the room have calibration data (e.g., a calibration procedure was previously performed in the room and calibration data is available in the media playback system), the method proceeds to block 850, where spatial calibration data is obtained from nearby devices in the room (or from another component of the media playback system). If, at block 848, the device is not in a room with existing calibration data, then, at block 852, the wearable playback device may directly determine calibration coefficients, for example, by performing a calibration process. Examples of suitable calibration processes can be found in commonly owned U.S. Patent No. 9,7906,323, entitled "Calibration of a Playback Device," and U.S. Patent No. 9,763,018, entitled "Calibration of an Audio Playback Device," each of which is incorporated herein by reference in its entirety. In various examples, the wearable playback device can have access to calibration data obtained by another playback device in a particular room or location, or the wearable playback device itself can perform a calibration procedure to determine the calibration coefficients.

[0132] If the wearable device itself performs or is otherwise involved in determining the room acoustics, the out-loud playback device can be configured to emit a calibration sound (e.g., a tone, sweep, chirp, or other suitable audio output) that can be detected and recorded via one or more microphones of the wearable playback device. The calibration audio data can then be used to perform a calibration (e.g., via the wearable playback device, the out-loud playback device, a control device, a remote computing device, or any other suitable device or combination of devices) to determine a room acoustic profile applicable to the wearable playback device output. Optionally, the calibration determination can also be used to calibrate the audio output of other devices in the room, such as, for example, the out-loud playback device employed to emit the calibration sound.

[0133] At block 854, based on the device type and calibration data, the method may determine wearable playback device audio parameters to resemble an acoustic device at the identified location. For example, using the calibration data from blocks 850 or 852, playback through the wearable playback device may be modified to resemble the acoustics of loud playback at the location. Once the audio parameters are determined, method 840 may proceed to method 800 shown in FIG. 8A.

[0134] FIG. 9 illustrates an example scenario for playing generated audio content via a wearable playback device within a home environment. As illustrated, a user wearing a wearable playback device can move to various locations 901-905 within the home environment. Optionally, the wearable playback device can wirelessly connect (e.g., via Bluetooth, WiFi, etc.) to nearby playback devices to relay data (e.g., audio content, playback control, timing data, and / or sensor data, etc.). In the illustrated scenario, at location 901, a user might be working out in their study and listening to an upbeat-tempo exercise soundscape. When the user moves to location 902 in their kitchen, the wearable playback device can automatically transition to playing more relaxing generated audio content suitable for cooking, eating, or socializing. When the user moves to location 903 in their home office, the wearable playback device can automatically transition to playing generated audio content suitable for work or concentration. Next, at location 904 in the television viewing room, the wearable playback device can transition to receiving audio data from the home theater primary playback device to play audio accompanying the video content being played through the television. Finally, when the user enters the bedroom at location 905, the wearable playback device can automatically transition from playing TV audio to outputting relaxing generated audio content to help the user relax in preparation for sleep. In these and other examples, the wearable playback device can automatically transition playback by switching from the generated audio content to other content sources or by modulating the generated audio content itself based on the user's location and / or other contextual data.

[0135] c. Playback control between wearable playback devices and out-of-home playback devices 10A shows an example method 1000 for swapping playback of generated audio content between wearable playback device(s) and out-loud playback device(s), and FIG. 10B shows an example scenario including swapping playback of generated audio content between wearable playback device(s) and out-loud playback device(s). Method 1000 begins with initiating a swap at block 1002. The swap may include moving (taking over) playback of the generated audio output from one device (e.g., the wearable playback device) to one or more nearby target devices (e.g., out-loud playback devices within substantial earshot of the wearable device).

[0136] In block 1006, in examples where the generated audio is swapped to one or more target devices, an outloud target group coordinator device can be selected from among the target devices. The group coordinator can be selected based on parameters such as network connectivity, processing power, available memory, remaining battery life, and location. In some examples, the outloud group coordinator is selected based on the predicted direction of the user's path through the environment. For example, in the example of FIG. 10B, if the user is traveling from the master bedroom (bottom right) through the living room toward the kitchen, a playback device in the kitchen may be selected as the group coordinator.

[0137] In block 1008, the group coordinator may then determine the playback responsibilities of the target device, which may depend on the device's location, audio output capabilities (e.g., subwoofer vs. ultraportable), and other suitable parameters.

[0138] Block 1010 maps the resulting audio content channels to the playback responsibilities determined in block 1008, and block 1012 transfers playback from the wearable playback device to the target out-loud playback device(s). In some embodiments, this swap involves gradually decreasing the playback volume on the wearable playback device and gradually increasing the playback volume on the out-loud playback device. Alternatively, this swap may involve abruptly stopping playback on the wearable playback device and simultaneously starting playback on the out-loud playback device.

[0139] There may be certain rules for selecting a group coordinator, such as not using portable and / or subwoofer devices, but in some instances it may make sense to use a subwoofer or similar device as a group coordinator because its low frequency output is available regardless of the user's location and is therefore more likely to remain in the soundscape as the user moves throughout the home.

[0140] In some examples, the group coordinator is not the target device but instead another local device (e.g., a local hub device) or a remote computing device such as a cloud server. In certain examples, the wearable device continues to generate generated audio content and coordinates the audio between the swapped target devices.

[0141] Figure 10C is a schematic diagram of a system 1020 for generating and playing generated media content according to aspects of the present disclosure. System 1020 is similar to systems 600 and 700 described above with respect to Figures 6 and 7, and certain components omitted in Figure 10C are included in various forms. Some portions of the generation of the generated media content have been omitted, and only the generated media module 211 and the resulting generated media content 616 are shown here. However, in various embodiments, any of the approaches or techniques described elsewhere herein or known to those skilled in the art can be incorporated into the generation of the generated media content 616.

[0142] In various examples, generated media content 616 can include multi-channel media content. The generated media content 616 is then sent to either wearable playback device 102o and / or group coordinator 1022 for playback via outloud playback devices 102a, 102b, and 102c. In the case of a swap action, the generated media content 616 being played on wearable playback device 102o can be swapped to group coordinator 1022 such that playback on wearable playback device 102o stops and playback on outloud playback devices 102a-102c begins. The reverse swap can also be performed, in which the outloud generated audio content stops playing on outloud playback devices 102a-102c and begins playing on wearable playback device 102o.

[0143] In some embodiments, wearable playback device 102o can include a generated media module 211 installed on it, and swapping playback from wearable playback device 102o to an out-of-home playback device includes transferring the generated audio content being generated on wearable playback device 102o from wearable playback device 102o to group coordinator 1022.

[0144] In some examples, swapping to the set of target out-loud playback devices involves sending details or parameters of the generated media module to the target group coordinator 1022. In other examples, the wearable device 102i receives the generated audio from a cloud-based generated media module, and swapping the audio from the wearable device 102o to the target devices 102a-102c involves the target group coordinator 1022 requesting the generated media content from the cloud-based generated media module 211.

[0145] Some or all of the multiple playback devices 102 can be configured to receive one or more input parameters 603. As mentioned above, the input parameters 603 can include any suitable input, such as user input (e.g., user selection via a controller device), sensor data (e.g., the number of people present in a room, background noise levels, time of day, weather data, etc.), or other suitable input parameters. In various examples, the input parameter(s) 603 can be optionally provided to the playback device 102 and / or detected or determined by the playback device 102 itself.

[0146] In some examples, to determine specific playback responsibilities and coordinate synchronized playback among various devices, the group coordinator 1022 can send timing information and / or playback responsibility information to the playback devices 102. Additionally or alternatively, the playback devices 102 themselves can determine their respective playback responsibilities based on the multi-channel media content received along with the input parameters 603.

[0147] 10D shows an example method 1030 for playing generated audio content through both a wearable playback device and an out-loud playback device. At block 1032, method 1030 plays generated audio from the out-loud playback device(s) at a first volume level, which is lower than a not-yet-reached ultimate second volume level, while continuing to play the generated audio from the wearable playback device. The wearable playback device can be tethered in transparent mode at block 1034, and the volume of the wearable playback device is gradually reduced from a fourth volume level to a third volume level less than the fourth volume level at block 1036. For example, the fourth volume level can be the current playback volume, and the third volume level can be an intermediate volume level between the fourth volume level and mute.

[0148] At block 1038, the playback volume of the out-loud playback device(s) is gradually increased from a first volume level to a second volume level that is greater than the first volume level, and playback of the generated audio content from the wearable playback device is completely stopped at block 1040. In various embodiments, the gradual decrease in playback volume of the wearable playback device may occur simultaneously with the gradual increase in playback volume of the out-loud playback device.

[0149] In some examples, while in transparency mode (block 1034), audio played from the wearable playback device is modified to reduce the amount of audio played in frequency bands associated with human speech. This allows speech to be heard even while sound is being played on the wearable playback device. Additionally or alternatively, active noise cancellation functionality can be reduced or turned off entirely during transparency mode.

[0150] d. Rules engine for restricting generated media playback In some cases, it may be desirable to impose user-level and / or location-level restrictions on the playback of soundscapes and other audio content. FIG. 11 illustrates an example rules engine for restricting the playback of generated audio content within an environment. As illustrated, different users may have different location-specific restrictions. In this example, User 1 is allowed to play audio content in some locations (patio, entry den, kitchen / dining room, living room, bedroom 1, bathroom 1) but is restricted from playing audio content in other locations (bedroom 2, bathroom 2). Similarly, User 2 is allowed to play audio content in some locations (patio, bedroom 2, entry den, kitchen / dining room, living room) but is restricted from playing audio content in other locations (bathroom 2, bedroom 1, bathroom 1). Such restrictions may be based on user preferences, device capabilities, or user characteristics (e.g., age) and may prohibit a user from playing soundscapes or other audio in certain zones or rooms within an environment or home.

[0151] IV. Conclusion The above description discloses various exemplary systems, methods, apparatus, and articles of manufacture that include, among other things, firmware and / or software executing on hardware. Such examples are merely illustrative and not limiting. For example, it is contemplated that any or all aspects or components of the firmware, hardware, and / or software may be embodied exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Accordingly, the provided examples are not the only ways to implement such systems, methods, apparatus, and / or articles of manufacture.

[0152] As used herein, sending information to a particular component, device, and / or system should be understood to include indirectly or directly sending information (e.g., a message, a request, a response) to the particular component, device, and / or system. Thus, information sent to a particular component, device, and / or system may pass through any number of intermediate components, devices, and / or systems before reaching its destination. For example, a control device may send information to a playback device by first sending the information to a computing system, which then sends the information to the playback device. Additionally, the information may be modified by an intermediary component, device, and / or system. For example, the intermediary component, device, and / or system may modify portions of the information, reformat the information, and / or incorporate additional information.

[0153] Similarly, as used herein, receiving information from a particular component, device, and / or system should be understood to include receiving information (e.g., a message, a request, a response) indirectly or directly from the particular component, device, and / or system. Thus, information received from a particular component, device, and / or system may pass through any number of intermediate components, devices, and / or systems before being received. For example, a control device may receive information indirectly from a playback device by receiving information originating from the playback device from a cloud server. Furthermore, information may be modified by an intermediary component, device, and / or system. For example, the intermediary component, device, and / or system may modify portions of the information, reformat the information, and / or incorporate additional information.

[0154] This specification is presented primarily in terms of exemplary environments, systems, procedures, steps, logic blocks, processes, and other symbolic representations that directly or indirectly resemble the operation of network-coupled data processing devices. Such process descriptions and representations are generally used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art. Numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, those skilled in the art will understand that certain embodiments of the present disclosure may be practiced without certain specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail to avoid unnecessarily obscuring aspects of the embodiments. Accordingly, the scope of the present disclosure is defined by the appended claims, rather than by the foregoing description of the embodiments.

[0155] The various examples described herein include one or more operations, functions, or actions represented by blocks. Although the blocks are shown in sequential order, these blocks may also be performed in parallel and / or in orders different from those disclosed and described herein. Additionally, various blocks may be combined into fewer blocks, divided into additional blocks, and / or eliminated based on the desired implementation.

[0156] Furthermore, for the methods disclosed herein, flowcharts illustrate the functions and operations of some example possible implementations. In this regard, each block may correspond to a module, segment, or portion of program code, including one or more instructions executable by one or more processors to implement specific logical functions or steps in the process. The program code may be stored on any type of computer-readable medium, such as a storage device including a disk or hard drive. The computer-readable medium may include non-transitory computer-readable media, such as register memory, processor cache, and tangible non-transitory computer-readable media for storing short-term data, such as random access memory (RAM). The computer-readable medium may also include non-transitory media, such as secondary or persistent long-term storage devices, such as read-only memory (ROM), optical or magnetic disks, and compact disk read-only memory (CD-ROM). The computer-readable medium may also be any other volatile or non-volatile storage system. The computer-readable medium may be considered, for example, a computer-readable storage medium or a tangible storage device. Furthermore, for the methods disclosed herein and other processes and methods, each block in the figures may represent circuitry hardwired to perform a particular logical function within the process.

[0157] If any of the appended claims are read to cover purely software and / or firmware embodiments, at least one of the elements in at least one example is expressly defined hereby to include a tangible, non-transitory medium, such as a memory, DVD, CD, Blu-ray®, etc., that stores the software and / or firmware.

[0158] V. Examples The disclosed technology is illustrated, for example, according to various examples described below. Various examples of embodiments of the disclosed technology are described as numbered examples for convenience. These are provided as examples and are not intended to limit the disclosed technology. It should be noted that any of the subordinate examples may be combined in any combination or may be included in their own independent examples. Other examples may be presented as well.

[0159] and one or more tangible, non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the wearable playback device to perform operations including: detecting that the wearable playback device is worn by a user; obtaining one or more input parameters via a network interface of the wearable playback device; generating generated media content based at least in part on the one or more input parameters after detecting that the wearable playback device is worn by the user; playing the generated media content via the one or more audio transducers; detecting that the wearable playback device is no longer worn by the user; and stopping playback of the generated media content via the one or more audio transducers after detecting that the wearable playback device is no longer worn by the user.

[0160] Example 2: The wearable playback device of any one of the preceding examples, wherein the one or more input parameters include location information, and as the location information changes over time, the generated media content changes based at least in part on the changing location information.

[0161] Example 3: The wearable playback device of any one of the preceding examples, wherein the location information includes ID information of a room or space within the user's environment.

[0162] Example 4: The wearable playback device of any one of the preceding examples, wherein the one or more input parameters include information identifying a user.

[0163] Example 5: The wearable playback device of any one of the preceding examples, wherein generating the generated media content includes (i) sending at least one of the one or more input parameters to a second playback device via a network interface of the wearable playback device, and (ii) receiving the generated media content generated via the second playback device via the network interface of the wearable playback device.

[0164] Example 6: The wearable playback device of any one of the preceding examples, wherein the one or more input parameters are comprised of one or more first input parameters, and generating the generated media content includes receiving at least one second input parameter from a network interface of the second playback device via a network interface of the wearable playback device.

[0165] Example 4: The wearable playback device of any one of the preceding examples, wherein the one or more input parameters comprise sound data obtained via a microphone of the wearable playback device.

[0166] Example 8: The wearable playback device of any one of the preceding examples, wherein the sound data constitutes a seed for a generative media content engine.

[0167] Example 9: The wearable playback device of any one of the preceding examples, wherein the operations further include adjusting (on / off, changing color, etc.) one or more light sources in the user's environment based at least in part on the one or more input parameters concurrently with playback of the generated media content.

[0168] Example 10: The wearable playback device of any one of the preceding examples, wherein the one or more input parameters include spatial calibration information related to a user's position, and playing the generated media content includes playing audio via the wearable playback device using the spatial calibration information.

[0169] Example 11: The wearable playback device of any one of the preceding examples, wherein the operations further include initiating synchronized playback of at least a portion of the generated media content via a second audio playback device while playing the generated media content via the wearable playback device.

[0170] Example 12: The wearable playback device of any one of the preceding examples, wherein the operation further includes lowering a playback volume of the wearable playback device after initiating synchronized playback of the generated media content via the second audio playback device.

[0171] Example 13: The wearable playback device of any one of the preceding examples, wherein initiating synchronized playback of the generated media content via the second audio playback device is based at least in part on a user instruction.

[0172] Example 14: The wearable playback device of any one of the preceding examples, wherein initiating synchronized playback of the generated media content via the second audio playback device is based at least in part on a proximity determination between the wearable playback device and the second audio playback device.

[0173] Example 15: The wearable playback device of any one of the preceding examples, further including: detecting connection of a wireless audio source while playing generated media content via the wearable playback device; and, after detecting connection of the wireless audio source, stopping playback of the generated media content and starting playback of audio data from the wireless audio source.

[0174] Example 16: The wearable playback device of any one of the preceding examples, further including: determining that the wireless audio source has been disconnected; and, after determining that the wireless audio source has been disconnected, starting playback of the generated media content.

[0175] Example 17: The wearable playback device of any one of the preceding examples, wherein the generated media content is based at least in part on audio data from a wireless audio source.

[0176] Example 18: The wearable playback device of any one of the preceding examples, wherein the one or more input parameters are physiological sensor data (e.g., biosensors, wearable sensors (heart rate, temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); wearable playback device capability data (e.g., number and type of transducers, output power); wearable playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with other playback devices); or user data (e.g., user ID, number of users present, user location, user history data, user preference data, user biometric data (heart rate, body temperature, respiration rate, brain activity, voice speech characteristics), user mood data).

[0177] Example 19: A method comprising: detecting that a wearable playback device is being worn by a user; obtaining one or more input parameters via a network interface of the wearable playback device; generating generated media content based at least in part on the one or more input parameters after detecting that the wearable playback device is being worn by the user; playing the generated media content via the wearable playback device; detecting that the wearable playback device is no longer being worn by the user; and stopping playback of the generated media content via the wearable playback device after detecting that the wearable playback device is no longer being worn by the user.

[0178] Example 20: The method of any one of the preceding examples, wherein the one or more input parameters include location information, and as the location information changes over time, the generated media content changes based at least in part on the changing location information.

[0179] Example 21: The method of any one of the preceding examples, wherein the location information includes identifying a room or space within the user's environment.

[0180] Example 22: The method of any one of the preceding examples, wherein the one or more input parameters include information identifying a user.

[0181] Example 23: The method of any one of the preceding examples, wherein the step of generating the generated media content includes (i) sending at least one of the one or more input parameters to a second playback device via a network interface of the wearable playback device, and (ii) receiving the generated media content generated via the second playback device via the network interface of the wearable playback device.

[0182] Example 24: The method of any one of the preceding examples, wherein the one or more input parameters are comprised of one or more first input parameters, and generating the generated media content includes receiving at least one second input parameter from a network interface of the second playback device via a network interface of the wearable playback device.

[0183] Example 25: The method of any one of the preceding examples, wherein the one or more input parameters comprise sound data obtained via a microphone of a wearable playback device.

[0184] Example 26: The method of any one of the preceding examples, wherein the sound data constitutes a seed for a generative media content engine.

[0185] Example 27: The method of any one of the preceding examples, further comprising adjusting (on / off, changing color, etc.) one or more light sources in the user's environment based at least in part on the one or more input parameters concurrently with playback of the generated media content.

[0186] Example 28: The method of any one of the preceding examples, wherein the one or more input parameters include spatial calibration information related to a user's position, and playing the generated media content includes playing audio via the wearable playback device using the spatial calibration information.

[0187] Example 29: The method of any one of the preceding examples, further comprising initiating synchronized playback of at least a portion of the generated media content via a second audio playback device while playing the generated media content via the wearable playback device.

[0188] Example 30: The method of any one of the preceding examples, further comprising: reducing the playback volume of the wearable playback device after initiating synchronized playback of the generated media content via the second audio playback device.

[0189] Example 31: The method of any one of the preceding examples, wherein initiating synchronized playback of the generated media content via the second audio playback device is based at least in part on a user instruction.

[0190] Example 32: The method of any one of the preceding examples, wherein initiating synchronized playback of the generated media content via the second audio playback device is based at least in part on a proximity determination between the wearable playback device and the second audio playback device.

[0191] Example 33: The method of any one of the preceding examples, further comprising detecting connection of a wireless audio source while playing generated media content via the wearable playback device, and, after detecting connection of the wireless audio source, stopping playback of the generated media content and starting playback of audio data from the wireless audio source.

[0192] Example 34: The method of any one of the preceding examples, further comprising: determining that the wireless audio source has been disconnected; and, after determining that the wireless audio source has been disconnected, starting playback of the generated media content.

[0193] Example 35: The method of any one of the preceding examples, wherein the generated media content is based at least in part on audio data from a wireless audio source.

[0194] Example 36: The method of any one of the preceding examples, wherein the one or more input parameters are physiological sensor data (e.g., biosensors, wearable sensors (heart rate, temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); wearable playback device capability data (e.g., number and type of transducers, output power); wearable playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with other playback devices); or user data (e.g., user ID, number of users present, user location, user history data, user preference data, user biometric data (heart rate, body temperature, respiration rate, brain activity, voice speech characteristics), user mood data).

[0195] Example 37: One or more tangible, non-transitory computer-readable media having stored thereon instructions for operations executed by one or more processors of a wearable playback device, the operations comprising: detecting that the wearable playback device is being worn by a user; obtaining one or more input parameters via a network interface of the wearable playback device; generating generated media content based at least in part on the one or more input parameters after detecting that the wearable playback device is being worn by the user; playing the generated media content via one or more audio transducers of the wearable playback device; detecting that the wearable playback device is no longer being worn by the user; and stopping playback of the generated media content via the wearable playback device after detecting that the wearable playback device is no longer being worn by the user.

[0196] Example 38: One or more computer-readable media described in any one of the preceding examples, wherein the one or more input parameters include location information, and as the location information changes over time, the generated media content changes based at least in part on the changing location information.

[0197] Example 39: One or more computer-readable media according to any one of the preceding examples, wherein the location information includes ID information of a room or space within the user's environment.

[0198] Example 40: One or more computer-readable media according to any one of the preceding examples, wherein the one or more input parameters include information identifying a user.

[0199] Example 41: One or more computer-readable media described in any one of the preceding examples, wherein generating the generated media content includes (i) sending at least one of the one or more input parameters to a second playback device via a network interface of the wearable playback device, and (ii) receiving the generated media content generated via the second playback device via the network interface of the wearable playback device.

[0200] Example 42: One or more computer-readable media described in any one of the preceding examples, wherein the one or more input parameters are comprised of one or more first input parameters, and generating the generated media content includes receiving at least one second input parameter from a network interface of the second playback device via a network interface of the wearable playback device.

[0201] Example 43: One or more computer-readable media according to any one of the preceding examples, wherein the one or more input parameters comprise sound data obtained via a microphone of a wearable playback device.

[0202] Example 44: One or more computer-readable media according to any one of the preceding examples, wherein the sound data constitutes a seed for a generative media content engine.

[0203] Example 45: One or more computer-readable media described in any one of the preceding examples, wherein the operations further include adjusting (on / off, changing color, etc.) one or more light sources in the user's environment based at least in part on the one or more input parameters simultaneously with playing the generated media content.

[0204] Example 46: One or more computer-readable media described in any one of the preceding examples, wherein the one or more input parameters include spatial calibration information related to a user's position, and playing the generated media content includes playing audio via the wearable playback device using the spatial calibration information.

[0205] Example 47: One or more computer-readable media described in any one of the preceding examples, wherein the operations further include initiating synchronized playback of at least a portion of the generated media content via a second audio playback device while playing the generated media content via the wearable playback device.

[0206] Example 48: One or more computer-readable media described in any one of the preceding examples, wherein the operation further includes lowering a playback volume of the wearable playback device after initiating synchronized playback of the generated media content via the second audio playback device.

[0207] Example 49: One or more computer-readable media described in any one of the preceding examples, wherein initiating synchronized playback of the generated media content via the second audio playback device is based at least in part on a user instruction.

[0208] Example 50: One or more computer-readable media described in any one of the preceding examples, wherein initiating synchronized playback of the generated media content via the second audio playback device is based at least in part on a proximity determination between the wearable playback device and the second audio playback device.

[0209] Example 51: One or more computer-readable media described in any one of the preceding examples, further including: detecting connection of a wireless audio source while playing generated media content via the wearable playback device; and, after detecting connection of the wireless audio source, stopping playback of the generated media content and starting playback of audio data from the wireless audio source.

[0210] Example 52: One or more computer-readable media described in any one of the preceding examples, further including: determining that the wireless audio source has been disconnected; and, after determining that the wireless audio source has been disconnected, starting playback of the generated media content.

[0211] Example 53: One or more computer-readable media of any one of the preceding examples, wherein the generated media content is based at least in part on audio data from a wireless audio source.

[0212] Example 54: One or more computer-readable media according to any one of the preceding examples, wherein the one or more input parameters are physiological sensor data (e.g., biosensors, wearable sensors (heart rate, temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); wearable playback device capability data (e.g., number and type of transducers, output power); wearable playback device status (e.g., device temperature, battery level, current audio playback, location of playback device, whether the playback device is combined with other playback devices); or user data (e.g., user ID, number of users present, user location, user history data, user preference data, user biometric data (heart rate, body temperature, respiration rate, brain activity, voice speech characteristics), user mood data).

Claims

1. 1. A method comprising: obtaining one or more input parameters via a network interface of the wearable playback device; generating generated media content based at least in part on the one or more input parameters; playing the generated media content via the wearable playback device after detecting that the wearable playback device is worn by the user; and stopping playback of the generated media content via the wearable playback device when the wearable playback device is no longer worn by the user; A method comprising:

2. The method of claim 1 , wherein the one or more input parameters include location information, and as the location information changes over time, the generated media content changes based at least in part on the changing location information.

3. The method of claim 2 , wherein the location information includes information identifying a room or space within the user's environment.

4. 10. A method according to any one of the preceding claims, wherein the one or more input parameters include information identifying a user.

5. 10. The method of claim 1, wherein the step of generating the generated media content comprises: (i) transmitting at least one of the one or more input parameters to a second playback device via a network interface of the wearable playback device; and (ii) receiving the generated media content generated via the second playback device via a network interface of the wearable playback device.

6. 10. The method of claim 1, wherein the one or more input parameters include one or more first input parameters, and wherein generating the generated media content includes receiving at least one second input parameter from a network interface of the second playback device via a network interface of the wearable playback device.

7. 10. The method of claim 1, wherein the one or more input parameters comprise sound data obtained via a microphone of a wearable playback device.

8. The method of claim 7 , wherein the sound data constitutes a seed for a generative media content engine.

9. 10. The method of claim 1, further comprising adjusting (turning on / off, changing color, etc.) one or more light sources in the user's environment based at least in part on the one or more input parameters concurrently with the playback of the generated media content.

10. 10. The method of claim 1, wherein the one or more input parameters include spatial calibration information related to a user's position, and wherein playing the generated media content includes playing audio via the wearable playback device using the spatial calibration information.

11. 10. The method of claim 1, further comprising: initiating synchronized playback of at least a portion of the generated media content via a second audio playback device while playing the generated media content via the wearable playback device.

12. 12. The method of claim 11, further comprising: reducing a playback volume of the wearable playback device after initiating synchronized playback of the generated media content via the second audio playback device.

13. 13. The method of claim 11, wherein initiating synchronized playback of the generated media content via the second audio playback device is based at least in part on a user instruction.

14. 14. The method of claim 11, wherein initiating synchronized playback of the generated media content via the second audio playback device is based at least in part on a proximity determination between the wearable playback device and the second audio playback device.

15. Furthermore, detecting a connection of a wireless audio source while playing generated media content via the wearable playback device; after detecting a connection of the wireless audio source, stopping the playback of the generated media content and starting the playback of audio data from the wireless audio source; 10. The method of any one of the preceding claims, comprising:

16. Furthermore, determining that the wireless audio source has been disconnected; after determining that the wireless audio source has been disconnected, commencing playback of the generated media content; The method of claim 15, comprising:

17. 17. The method of claim 15 or 16, wherein the generated media content is based at least in part on audio data from a wireless audio source.

18. The one or more input parameters are: physiological sensor data, network device sensor data, environmental data, Wearable playback device capability data, Wearable playback device state, or User Data 10. The method of any one of the preceding claims, comprising one or more of:

19. One or more tangible, non-transitory computer-readable media having stored thereon instructions for operations to be performed by one or more processors of a wearable playback device according to the method of any one of the preceding claims.

20. A wearable playback device, one or more audio transducers; Network interface; one or more processors; and One or more tangible, non-transitory computer-readable media storing instructions for operations to be performed by one or more processors of a wearable playback device according to the method of any one of the preceding claims. A wearable playback device equipped with

Citation Information

Patent Citations

  • Portable audio equipment and on-vehicle audio equipment

    JP2002373484A

  • Headphone device

    JP2003037886A

  • Headset and headset power management

    JP2007104670A

  • Reproduction device, reproduction method, program, and reproduction system

    WO2018055959A1

  • Playback of generative media content

    WO2022109556A2