Playback of generative media content

A distributed architecture for media playback systems addresses synchronization challenges by using remote computing for content generation and local synchronization, ensuring real-time, dynamic media playback across multiple devices.

JP2025179181APending Publication Date: 2025-12-09SONOS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025147155
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-30
Filing Date
2025-09-04
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Coordinating the playback of dynamically generated media content across multiple devices in an environment is challenging due to computational limitations and the need for real-time synchronization without latency.

Method used

A distributed architecture is employed, where remote computing devices generate permutations of media content and local playback devices receive and synchronize playback based on input parameters, with a coordinator device managing channel distribution among devices.

Benefits of technology

Enables synchronized, real-time playback of dynamically generated media content across multiple devices without latency, allowing for dynamic adjustments based on environmental inputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025179181000001_ABST
    Figure 2025179181000001_ABST
Patent Text Reader

Abstract

To provide a method for simultaneously playing generated media content (e.g., generated audio) across multiple playback devices, and a coordinator device.SOLUTION: A method includes the steps of: receiving input parameters at a coordinator device that facilitates simultaneous playback by routing media content, associated data, and / or instructions to member devices; transmitting the input parameters from the coordinator device to multiple playback devices, each having a generated media module internally; sending timing data from the coordinator device to the multiple playback devices so that the playback devices simultaneously play back the generated media content at least partially based on the input parameters; and simultaneously playing back the generated media content via the playback devices at least partially based on the input parameters.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 198,866, filed November 18, 2020, entitled "Multi-Device Playback of Generated Media Content," U.S. Application No. 17 / 302,690, filed May 10, 2021, entitled "Playback of Generated Media Content," and U.S. Provisional Application No. 63 / 261,893, filed September 30, 2021, entitled "Multi-Channel Playback of Generated Media Content," each of which is incorporated by reference in its entirety.

[0002] The present disclosure relates to consumer goods, and more particularly to methods, systems, products, features, services, and other elements directed to media playback or some aspect thereof. [Background technology]

[0003] Options for accessing and listening to digital audio at high volume settings were limited until 2002, when SONOS, Inc. began developing a new type of playback system. Sonos then filed one of the first patent applications, titled "Method for Synchronizing Audio Playback between Multiple Networked Devices," in 2003 and began offering its first commercially available media playback system in 2005. The Sonos wireless home sound system allows people to experience music from many sources through one or more networked playback devices. Through a software control application installed on a controller (e.g., smartphone, tablet, computer, audio input device), users can play what they want in any room with a networked playback device. Media content (e.g., songs, podcasts, video sounds) can be streamed to the playback devices so that each room with a playback device can play corresponding different media content. Furthermore, rooms can be grouped together for synchronized playback of the same media content and / or the same media content can be listened to synchronously in all rooms. [Brief explanation of the drawings]

[0004] The features, aspects, and advantages of the disclosed technology may be better understood with regard to the following description, appended claims, and accompanying drawings, set forth below. Those skilled in the art will recognize that the features shown in the drawings are for illustrative purposes and that variations, including different and / or additional features and arrangements thereof, are possible. [Figure 1A] 1 is a partial cutaway view of an environment having a media playback system configured in accordance with aspects of the disclosed technology. [Figure 1B] 1B is a schematic diagram of the media playback system and one or more networks of FIG. [Figure 1C] Playback device block diagram [Figure 1D] Playback device block diagram [Figure 1E] Block diagram of a combined playback device [Figure 1F] Network Microphone Device Block Diagram [Figure 1G] Playback device block diagram [Figure 1H] Partial schematic diagram of the control device [Figure 1I] Schematic diagram of supported media playback system zones [Figure 1J] Schematic diagram of supported media playback system zones [Figure 1K] Schematic diagram of supported media playback system zones [Figure 1L] Schematic diagram of supported media playback system zones [Figure 1M] Media Playback System Area Schematic [Figure 2] 1 is a functional block diagram of a system for playback of generated media content according to an example of the present technology; [Figure 3] FIG. 1 is a functional block diagram of a generated media module according to aspects of the present technology. [Figure 4] FIG. 1 illustrates an example architecture for storing and retrieving generated media content in accordance with aspects of the present technology. [Figure 5] FIG. 1 is a functional block diagram illustrating data exchange in a system for playback of generated media content in accordance with aspects of the present technology. [Figure 6] FIG. 1 is a schematic diagram of an exemplary distributed generative media playback system in accordance with aspects of the present technology. [Figure 7] Diagram of a generative media playback system for multi-channel playback [Figure 8] Diagram of another generative media playback system for multi-channel playback [Figure 9] 1 is a flow diagram of a method for playing generated media content in accordance with aspects of the present technology; [Figure 10] 1 is a flow diagram of a method for playing generated media content in accordance with aspects of the present technology; [Figure 11] 1 is a flow diagram of a method for playing generated media content in accordance with aspects of the present technology; [Figure 12] 1 is a flow diagram of a method for playing generated media content in accordance with aspects of the present technology; [Figure 13] 1 is a flow diagram of a method for playing generated media content in accordance with aspects of the present technology;

[0005] The drawings are for purposes of illustrating examples of the technology; however, one skilled in the art would understand that the technology disclosed herein is not limited to the arrangements and / or instrumentality shown in the drawings. DETAILED DESCRIPTION OF THE INVENTION

[0006] I. Overview Generative media content is content that is dynamically synthesized, created, and / or modified based on algorithms, whether implemented in software or physical models. Generative media content can change over time based solely on algorithms or in conjunction with contextual data (e.g., user sensor data, environmental sensor data, occurrence data). In various examples, such generative media content can include generative audio (e.g., music, ambient sounds, etc.), generative visual images (e.g., abstract visual designs that dynamically change shape, color, etc.), or any other suitable media content or combination thereof. As described elsewhere herein, generative audio can be generated, at least in part, via algorithms and / or non-human systems that utilize rule-based computations to generate novel audio content.

[0007] Because generated media content can change dynamically in real time, it enables unique user experiences not available using traditional media playback of pre-recorded content. For example, the generated audio can be endless and / or dynamic audio that changes as inputs to the algorithm (e.g., input parameters related to user input, sensor data, media source data, or any other suitable input data) change. In some examples, the generated audio can be used to steer a user's mood toward a desired emotional state, with one or more characteristics of the generated audio changing in response to real-time measurements that reflect the user's emotional state. As used in examples of the present technology, a system can provide generated audio based on the user's current and / or desired emotional state, based on the user's activity level, based on the number of users present in the environment, or based on any other suitable input parameters.

[0008] As another example, the generated audio can be created and / or modified based on one or more inputs, such as a user's location or activity, the number of users present in the room, the time of day, or any other input (e.g., as determined by one or more sensors or user input). For example, when one user is sitting at their desk in a calm state, a media playback system can automatically generate generated audio content suitable for intensive study or work, while when multiple users are present in the room in an excited state with a lot of movement, the same media playback system can automatically generate generated audio suitable for a social gathering or dance party. In various examples, audio characteristics that can be dynamically modified to produce the generated audio can include audio sample or clip selection, tempo, bass / treble / mid-range volume, spatial filtering of the audio output, or any other suitable audio characteristics. Audio characteristics can be modified by using audio samples that may have different tones or sounds, timing of tones or sounds, and / or desired qualities. In some cases, the playback of content can also be modified by filtering or modulating characteristics, such as equalization, phase, or reverb / delay. During the listening experience, the audio characteristics of the generated music can be altered based on several inputs, such as the time of day, geographic location, weather, or various user inputs, such as inferred mood, collective level of activity, or physiological inputs such as heart rate.

[0009] In an environment including multiple individual playback devices, coordinating the playback of generated audio content across the various playback devices can be difficult. In some cases, each playback device can synchronize and play the same generated audio content. To do so, the various devices can synchronize both their inputs or other parameters for the generated media content modules and the playback of the resulting generated audio. In some examples, some or all of the playback devices can have different playback responsibilities from one another (e.g., corresponding to different channels of audio input or other such division of playback responsibilities), but playback can still occur simultaneously (e.g., synchronously) for listening by one or more users in the environment. In some examples, the different playback devices can play entirely separate generated audio content that can nevertheless be played simultaneously and / or synchronously. For example, in a room with a jungle-like visual décor, a first playback device can play generated audio corresponding to the sound of running water to simulate a stream, a second playback device can play generated audio corresponding to bird songs or other animal noises, and a third playback device can play generated audio corresponding to a rhythmic beat. Although each playback device outputs independently generated audio content, the user experience is nevertheless improved by all three devices simultaneously playing their respective generated audio content.

[0010] In these and other examples, it may be useful to coordinate playback among various playback devices. In some examples, a generated media group may include multiple individual devices that, during operation, play generated audio content simultaneously with each other. One device in the group may function as a coordinator device, while the remaining group devices function as member devices responsible for playback. During operation, the coordinator device may route media content, associated data, and / or instructions to the member devices to facilitate simultaneous playback. In some examples, the coordinator device includes a generated media module that may generate one or more streams of generated audio content based on one or more inputs (e.g., sensor data, user input, selected audio content sources, etc.). The generated audio content streams may then be transmitted to group member devices for playback. In some examples, the coordinator device itself may also become a member device, for example, by participating in audio playback.

[0011] Additionally or alternatively, one or more member devices may utilize their own generated media modules to dynamically generate generated audio content based on one or more input parameters. In such cases, the coordinator device may send instructions, data (e.g., timing data to facilitate synchronized playback), and / or input parameters to the member devices, which may then generate generated audio for playback in real time or near real time simultaneously with other devices in the group. Further examples are described in more detail below.

[0012] In some cases, processing input parameters to generate generated media content can be computationally intensive and may exceed the computational capabilities (e.g., processing power, available memory, etc.) of one or more local playback devices in an environment. Therefore, it may be useful to utilize a distributed architecture for generated media playback in which certain tasks required to generate generated media content are handled by a remote computing device (e.g., a cloud-based server) and other tasks are handled by one or more local playback devices. As an example, various permutations of generated media content may be generated and stored by one or more remote computing devices. These permutations may correspond to different energy levels, desired mood states, etc., and may be updated at the remote computing device over time. The local playback device may then query the remote computing device to receive a particular permutation of generated media content for playback. The particular permutation requested or delivered may be based, at least in part, on one or more input parameters that may be detected and / or provided by the playback device. In one example, a local playback device (or multiple such devices) may receive input parameters (e.g., sensor data) indicative of a number of people in a room. These parameters can indicate a high energy level, so the local playback device can request an appropriate permutation of the generated media content from the remote computing device, which can then select the appropriate permutation of the generated media content and send it to the local playback device for playback.

[0013] At the remote computing device(s), various permutations of generated media content can be generated and stored, each with different characteristics and / or profiles. For example, a generative media module stored on a remote computing device can generate multiple different permutations of generated media content using a particular generated media content model (e.g., an algorithm or set of rules that uses one or more audio segments and / or input parameters as input to generate new generated media content). For example, the generative media module can generate high-energy, medium-energy, and low-energy variations of the same generated media content, where the same (or at least some overlapping) audio segments are used across the various permutations, but the segments are mixed and / or modified differently to generate different content (e.g., higher or slower tempos, more or fewer chord changes, etc.).

[0014] Additionally or alternatively, several distinct audio segments may be stored locally on one or more playback devices within the local environment. These audio segments may be arranged, ordered, overlapped, mixed, and / or processed for playback in a manner that generates the generated media content. In some examples, a remote computing device may periodically provide instructions to the local playback device in the form of an updated generated media content model (e.g., an algorithm), which the local playback device may then use to play the locally stored distinct audio segments in a manner that achieves the desired psychoacoustic effect. In this example, the tasks required to output the generated audio are distributed such that the local playback device stores, arranges, and plays the constituent audio segments, while the remote computing device processes input parameters and determines how particular segments should be arranged and processed to generate the desired generated media content. Various other distributions of tasks between the local and remote computing devices are possible.

[0015] Multi-channel playback of generated media content can present particular challenges, especially considering the importance of synchronizing the playback of various channels across different playback devices in an environment. For example, in some cases, the specific distribution of generated media content among different playback devices may be changed in real time based on specific inputs (e.g., sensor data, user input, or other contextual information). While it may be useful to generate generated media content via a cloud server or other remote computing device, requiring such remote computing device to recalculate the channel distribution based on local context can result in undesirable latency.

[0016] The present technology addresses these and other problems by providing all channels of multi-channel generated media content (e.g., multi-channel content that includes at least some generated media content) to each of multiple playback devices in an environment. In some cases, this involves sending channels to a coordinator device, which then sends the channels to the playback devices in the environment. Each playback device may then receive instructions regarding a subset of channels (and at what levels) to play in synchronization with the other playback devices. For example, a playback device in a first area of ​​a room may play the sound of rain, while a playback device in another part of the room may play an accompanying rhythmic beat. As another example, each device may play two or more channels at different relative levels (e.g., a first playback device may play the sound of rain at 80% gain and an accompanying beat at 20% gain, while a second playback device does the opposite). These playback responsibilities and distribution of channels may change in real time based on one or more inputs. For example, as more users enter the room, the tempo of the beat may be increased or the relative levels of the various channels may be adjusted. By distributing all channels to all playback devices, such dynamic variations can be implemented quickly without delay attendant routing information back to a cloud-based server for updated calculations. In various examples, the specific playback responsibilities assigned to each device can be determined via a coordinator device, via a control device (e.g., a smartphone application or other component), via the playback device itself, or in other ways (e.g., a remote computing device can include metadata accompanying the multi-channel media content that indicates a default or recommended distribution of playback responsibilities).

[0017] Although some examples described herein may refer to functions performed by given parties, such as "users," "listeners," and / or other entities, it should be understood that this is for illustrative purposes only. The claims should not be construed as requiring action by any such example actors unless expressly required by the language of the claims themselves.

[0018] In the figures, like reference numbers generally indicate similar and / or identical elements. To facilitate the description of any particular element, the most significant digit(s) of the reference number refers to the figure in which that element is first introduced. For example, element 110a is first introduced and described with reference to FIG. 1A. Many of the details, dimensions, angles, and other features shown in the figures are merely illustrative of particular examples of the disclosed technology. Thus, other examples can have other details, dimensions, angles, and features without departing from the spirit or scope of the present disclosure. Moreover, those skilled in the art will recognize that further examples of the various disclosed technologies can be practiced without some of the details described below.

[0019] II. Appropriate Operating Environment 1A is a partial cutaway view of a media playback system 100 distributed within an environment 101 (e.g., a home). Media playback system 100 includes one or more playback devices 110 (individually identified as playback devices 110a-110n), one or more network microphone devices (“NMDs”) 120 (individually identified as NMDs 120a-120c), and one or more control devices 130 (individually identified as control devices 130a and 130b).

[0020] As used herein, the term "playback device" may generally refer to a network device configured to receive, process, and / or output data for a media playback system. For example, a playback device may be a network device that receives and processes audio content. In some examples, a playback device includes one or more transducers or speakers powered by one or more amplifiers. However, in other examples, a playback device includes one or neither of a speaker and an amplifier. For example, a playback device may include one or more amplifiers configured to drive one or more speakers external to the playback device via corresponding wires or cables.

[0021] Additionally, as used herein, the term NMD (i.e., "network microphone device") may generally refer to a network device configured for audio detection. In some examples, an NMD is a standalone device configured primarily for audio detection. In other examples, an NMD is incorporated into a playback device (or vice versa).

[0022] The term “control device” may generally refer to a network device configured to perform functions related to facilitating user access, control, and / or configuration of media playback system 100 .

[0023] Each of the playback devices 110 is configured to receive audio signals or data from one or more media sources (e.g., one or more remote servers or one or more local devices) and play the received audio signals or data as sound. One or more NMDs 120 are configured to receive voice word commands, and one or more control devices 130 are configured to receive user input. In response to the received spoken word commands and / or user input, the media playback system 100 can play audio via one or more of the playback devices 110. In certain examples, the playback devices 110 are configured to initiate playback of media content in response to a trigger. For example, one or more of the playback devices 110 can be configured to play a morning playlist upon detection of an associated trigger condition (e.g., a user's presence in the kitchen, detection of operation of the coffee machine). In some examples, for example, the media playback system 100 is configured to play audio from a first playback device (e.g., playback device 110a) in synchronization with a second playback device (e.g., playback device 110b). Interactions between playback device 110, NMD 120, and / or control device 130 of media playback system 100 configured according to various examples of the present disclosure are described in more detail below in conjunction with FIGS. 1B-1H.

[0024] 1A , environment 101 comprises a home with several rooms, spaces, and / or playback zones, including (clockwise from top left) master bathroom 101a, master bedroom 101b, second bedroom 101c, family room or den 101d, office 101e, living room 101f, dining room 101g, kitchen 101h, and outdoor patio 101i. While specific examples and examples are described below in the context of a home environment, the techniques described herein may be implemented in other types of environments. In some examples, for example, media playback system 100 may be implemented in one or more commercial settings (e.g., a restaurant, mall, airport, hotel, retail store, or other establishment), one or more vehicles (e.g., a sport utility vehicle, bus, car, watercraft, boat, airplane), multiple environments (e.g., a combination of home and vehicle environments), and / or other suitable environments where multi-zone audio may be desirable.

[0025] Media playback system 100 can include one or more playback zones, some of which may correspond to rooms within environment 101. Media playback system 100 can be established with one or more playback zones, after which additional zones can be added or removed to form the configuration shown in FIG. 1A, for example. Each zone can be named according to a different room or space, such as office 101e, master bathroom 101a, master bedroom 101b, second bedroom 101c, kitchen 101h, dining room 101g, living room 101f, and / or outdoor patio 101i. In some aspects, a single playback zone can include multiple rooms or spaces. In certain aspects, a single room or space can include multiple playback zones.

[0026] In the illustrated example of FIG. 1A , the master bathroom 101a, the second bedroom 101c, the office 101e, the living room 101f, the dining room 101g, the kitchen 101h, and the outdoor patio 101i each include one playback device 110, while the master bedroom 101b and the private room 101d include multiple playback devices 110. In the master bedroom 101b, the playback devices 110l and 110m may be configured to play audio content synchronously, for example, as individual ones of the playback devices 110, as a combined playback zone, as an integrated playback device, and / or any combination thereof. Similarly, in the private room 101d, the playback devices 110h-110j may be configured to play audio content synchronously, for example, as individual ones of the playback devices 110, as one or more combined playback devices, and / or as one or more integrated playback devices. Further details regarding combined and integrated playback devices are described below with respect to FIGS. 1B and 1E .

[0027] In some aspects, one or more of the playback zones in environment 101 may each be playing different audio content. For example, a user may be grilling on patio 101i and listening to hip hop music being played by playback device 110c, while another user may be preparing food in kitchen 101h and listening to classical music being played by playback device 110b. In another example, a playback zone may play the same audio content in sync with another playback zone. For example, a user may be in office 101e listening to playback device 110f playing the same hip hop music being played by playback device 110c on patio 101i. In some aspects, playback devices 110c and 110f play the hip hop music in sync so that the user perceives that the audio content is being played seamlessly (or at least substantially seamlessly) while moving between the different playback zones. Further details regarding audio playback synchronization between playback devices and / or zones can be found, for example, in U.S. Pat. No. 8,234,395, entitled "System and Method for Synchronizing Operation Among Multiple Independently Clocked Digital Data Processing Devices," which is incorporated herein by reference in its entirety.

[0028] a. Suitable media playback system 1B is a schematic diagram of media playback system 100 and cloud network 102. For ease of illustration, certain devices of media playback system 100 and cloud network 102 have been omitted from FIG. 1B. One or more communication links 103 (hereinafter "links 103") communicatively couple media playback system 100 and cloud network 102.

[0029] Link 103 may comprise, for example, one or more wired networks, one or more wireless networks, one or more wide area networks (WANs), one or more local area networks (LANs), one or more personal area networks (PANs), one or more telecommunications networks (e.g., one or more Global System for Mobile Communications (GSM) networks, Code Division Multiple Access (CDMA) networks, Long Term Evolution (LTE) networks, 5G communications networks, and / or other suitable data transmission protocol networks), etc. Cloud network 102 is configured to deliver media content (e.g., audio content, video content, photos, social media content) to media playback system 100 in response to requests transmitted from media playback system 100 via link 103. In some examples, cloud network 102 is further configured to receive data (e.g., voice input data) from media playback system 100 and, in response, transmit commands and / or media content to media playback system 100.

[0030] Cloud network 102 comprises computing devices 106 (separately identified as first computing device 106a, second computing device 106b, and third computing device 106c). Computing devices 106 may comprise individual computers or servers, such as, for example, media streaming service servers that store audio and / or other media content, voice service servers, social media servers, media playback system control servers, etc. In some examples, one or more of computing devices 106 comprise modules of a single computer or server. In particular examples, one or more of computing devices 106 comprise one or more modules, computers, and / or servers. Furthermore, while cloud network 102 is described above in the context of a single cloud network, in some examples, cloud network 102 comprises multiple cloud networks including communicatively coupled computing devices. Furthermore, while cloud network 102 is illustrated in FIG. 1B as having three of computing devices 106, in some examples, cloud network 102 comprises fewer (or more) than three computing devices 106.

[0031] Media playback system 100 is configured to receive media content from network 102 via link 103. The received media content may comprise, for example, a uniform resource identifier (URI) and / or a uniform resource locator (URL). For example, in some examples, media playback system 100 may stream, download, or retrieve data from a URI or URL corresponding to the received media content. Network 104 communicatively couples link 103 to at least some of the devices of media playback system 100 (e.g., one or more of playback device 110, NMD 120, and / or control device 130). Network 104 may include, for example, a wireless network (e.g., a WiFi network, Bluetooth, Z-Wave network, ZigBee, and / or other suitable wireless communication protocol network) and / or a wired network (e.g., a network including Ethernet, Universal Serial Bus (USB), and / or another suitable wired communication protocol). As will be appreciated by those skilled in the art, as used herein, "WiFi" can refer to several different communication protocols including, for example, Institute of Electrical and Electronics Engineers (IEEE) 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.11ac, 802.11ad, 802.11af, 802.11ah, 802.11ai, 802.11aj, 802.11aq, 802.11ax, 802.11ay, 802.15, etc., transmitted at 2.4 gigahertz (GHz), 5 GHz, and / or other suitable frequencies.

[0032] In some examples, network 104 comprises a dedicated communications network that media playback system 100 uses to send messages between individual devices and / or to transmit media content to and from media content sources (e.g., one or more of computing devices 106). In particular examples, network 104 is configured to be accessible only to devices within media playback system 100, thereby reducing interference and contention with other home devices. However, in other examples, network 104 comprises an existing home communications network (e.g., a home WiFi network). In some examples, link 103 and network 104 comprise one or more of the same networks. In some aspects, for example, link 103 and network 104 comprise a telecommunications network (e.g., an LTE network, a 5G network). Furthermore, in some examples, media playback system 100 is implemented without network 104, and devices comprising media playback system 100 can communicate with each other via, for example, one or more direct connections, PANs, telecommunications networks, and / or other suitable communications links.

[0033] In some examples, audio content sources may be periodically added or removed from media playback system 100. In some examples, for example, media playback system 100 performs media item indexing when one or more media content sources are updated, added, and / or removed from media playback system 100. Media playback system 100 may scan for identifiable media items in some or all folders and / or directories accessible to playback device 110 and generate or update a media content database comprising metadata (e.g., title, artist, album, track length) and other associated information (e.g., URI, URL) for each identifiable media item found. In some examples, for example, the media content database is stored on one or more of playback device 110, NMD 120, and / or control device 130.

[0034] In the illustrated example of FIG. 1B , playback devices 110l and 110m comprise group 107a. Playback devices 110l and 110m may be located in different rooms within a home and may be temporarily or permanently grouped together in group 107a based on user input received at control device 130a and / or another control device 130 within media playback system 100. Once arranged in group 107a, playback devices 110l and 110m may be configured to synchronously play the same or similar audio content from one or more audio content sources. In certain examples, for example, group 107a may include a combined zone in which playback devices 110l and 110m each include a left audio channel and a right audio channel of multi-channel audio content, thereby creating or enhancing a stereo effect of the audio content. In some examples, group 107a may include additional playback devices 110. However, in other examples, media playback system 100 omits group 107a and / or other grouped arrangements of playback devices 110.

[0035] Media playback system 100 includes NMDs 120a and 120d, each comprising one or more microphones configured to receive voice utterances from a user. In the illustrated example of FIG. 1B , NMD 120a is a standalone device, and NMD 120d is incorporated into playback device 110n. NMD 120a is configured to receive voice input 121, for example, from user 123. In some examples, NMD 120a transmits data related to the received voice input 121 to a voice assistant service (VAS) configured to (i) process the received voice input data and (ii) transmit corresponding commands to media playback system 100. In some aspects, for example, computing device 106c comprises one or more modules and / or servers of a VAS (e.g., SONOS®, AMAZON®, GOOGLE®, APPLE®, MICROSOFT®). Computing device 106c can receive the voice input data from NMD 120a via network 104 and link 103. In response to receiving the audio input data, computing device 106c processes the audio input data (i.e., "Play Hey Jude by the Beatles") and determines that the processed audio input includes a command to play a song (e.g., "Hey Jude"). Accordingly, computing device 106c transmits a command to media playback system 100 from one or more appropriate media services of playback devices 110 (e.g., via one or more of computing devices 106).

[0036] b. Suitable playback device 1C is a block diagram of a playback device 110a including an input / output 111. The input / output 111 can include an analog I / O 111a (e.g., one or more wires, cables, and / or other suitable communication links configured to carry analog signals) and / or a digital I / O 111b (e.g., one or more wires, cables, or other suitable communication links configured to carry digital signals). In some examples, the analog I / O 111a is an audio line-in connection, including, for example, an auto-sensing 3.5mm audio line-in connection. In some examples, the digital I / O 111b includes a Sony / Philips Digital Interface Format (S / PDIF) communication interface and / or cable and / or a Toshiba Link (TOSLINK) cable. In some examples, the digital I / O 111b includes a High-Definition Multimedia Interface (HDMI®) interface and / or cable. In some examples, Digital I / O 111b includes one or more wireless communication links, including, for example, radio frequency (RF), infrared, WiFi, Bluetooth, or another suitable communication protocol. In particular examples, Analog I / O 111a and Digital 111b include interfaces (e.g., ports, plugs, jacks) configured to accept connectors of cables that transmit analog and digital signals, respectively, without necessarily including cables.

[0037] Playback device 110a can receive media content (e.g., audio content including music and / or other sounds) from local audio source 105, for example, via input / output 111 (e.g., cable, wire, PAN, Bluetooth connection, ad-hoc wired or wireless communication network, and / or another suitable communication link). Local audio source 105 can comprise, for example, a mobile device (e.g., a smartphone, a tablet, a laptop computer) or another suitable audio component (e.g., a television, a desktop computer, an amplifier, a phonograph, a Blu-ray player, memory for storing digital media files). In some aspects, local audio source 105 includes a local music library on a smartphone, a computer, a network-attached storage (NAS), and / or another suitable device configured to store media files. In certain examples, one or more of playback device 110, NMD 120, and / or control device 130 comprise local audio source 105. However, in other examples, the media playback system omits local audio source 105 entirely. In some examples, playback device 110 a does not include input / output 111 and receives all audio content over network 104 .

[0038] Playback device 110a further comprises electronics 112, a user interface 113 (e.g., one or more buttons, knobs, dials, touch-sensitive surfaces, displays, touchscreens), and one or more transducers 114 (hereinafter referred to as “transducers 114”). Electronics 112 is configured to receive audio from an audio source (e.g., local audio source 105) via input / output 111, one or more of computing devices 106a-106c via network 104 (FIG. 1B), amplify the received audio, and output the amplified audio for playback via one or more of transducers 114. In some examples, playback device 110a optionally includes one or more microphones 115 (e.g., a single microphone, multiple microphones, a microphone array) (hereinafter referred to as “microphones 115”). In particular examples, for example, playback device 110a having one or more of optional microphones 115 may operate as an NMD configured to receive audio input from a user and correspondingly perform one or more actions based on the received audio input.

[0039] 1C , electronic device 112 includes one or more processors 112a (hereinafter referred to as “processor 112a”), memory 112b, software components 112c, network interface 112d, one or more audio processing components 112g (hereinafter referred to as “audio components 112g”), one or more audio amplifiers 112h (hereinafter referred to as “amplifiers 112h”), and power source 112i (e.g., one or more power supplies, power cables, power receptacles, batteries, induction coils, Power over Ethernet (POE) interfaces, and / or other suitable power sources). In some embodiments, electronic device 112 optionally includes one or more other components 112j (e.g., one or more sensors, a video display, a touch screen, a battery charging base).

[0040] The processor 112a may comprise a clocked computing component configured to process data, and the memory 112b may comprise a computer-readable medium (e.g., a tangible, non-transitory computer-readable medium, data storage device) configured to store instructions for performing various operations and / or functions. The processor 112a is configured to execute the instructions stored in the memory 112b to perform one or more of the operations. The operations may include, for example, causing the playback device 110a to retrieve audio data from an audio source (e.g., one or more of the computing devices 106a-106c (FIG. 1B)) and / or another one of the playback devices 110. In some examples, the operations further include causing the playback device 110a to transmit the audio data to another one of the playback devices 110a and / or to another device (e.g., one of the NMDs 120). Particular examples include pairing the playback device 110a with another of the one or more playback devices 110 to enable a multi-channel audio environment (e.g., stereo pair, combined zone).

[0041] The processor 112a may be further configured to perform operations that cause the playback device 110a to synchronize playback of the audio content with another of the one or more playback devices 110. As will be appreciated by those skilled in the art, during synchronized playback of audio content on multiple playback devices, a listener preferably cannot perceive a time delay difference between the playback of the audio content by the playback device 110a and the playback of the audio content by one or more other playback devices 110. Further details regarding audio playback synchronization between playback devices may be found, for example, in U.S. Patent No. 8,234,395, incorporated by reference above.

[0042] In some examples, memory 112b is further configured to store data associated with playback device 110a, such as one or more zones and / or zone groups of which playback device 110a is a member, audio sources accessible to playback device 110a, and / or playback queues to which playback device 110a (and / or other playback devices of the one or more playback devices) may be associated. The stored data may include one or more state variables that are periodically updated and used to describe the state of playback device 110a. Memory 112b may also include data associated with the state of one or more of the other devices of media playback system 100 (e.g., playback device 110, NMD 120, control device 130). In some aspects, for example, state data is shared among at least some of the devices of media playback system 100 at predetermined time intervals (e.g., every 5 seconds, every 10 seconds, every 60 seconds), so that one or more of the devices have up-to-date data associated with media playback system 100.

[0043] Network interface 112d is configured to facilitate the transmission of data between playback device 110a and one or more other devices on a data network, such as link 103 and / or network 104 (FIG. 1B). Network interface 112d is configured to send and receive data corresponding to media content (e.g., audio content, video content, text, photographs) and other signals (e.g., non-transient signals) including digital packet data that include an Internet Protocol (IP)-based source address and / or an IP-based destination address. Network interface 112d can parse the digital packet data so that electronic device 112 appropriately receives and processes the data destined for playback device 110a.

[0044] In the depicted example of FIG. 1C , the network interface 112d includes one or more wireless interfaces 112e (hereinafter referred to as “wireless interface 112e”). The wireless interface 112e (e.g., a suitable interface including one or more antennas) can be configured to wirelessly communicate with one or more other devices (e.g., one or more of the other playback devices 110, the NMD 120, and / or the control device 130) communicatively coupled to the network 104 ( FIG. 1B ) according to a suitable wireless communication protocol (e.g., WiFi, Bluetooth, LTE). In some examples, the network interface 112d optionally includes a wired interface 112f (e.g., an interface or receptacle configured to receive a network cable, such as an Ethernet, USB-A, USB-C, and / or Thunderbolt cable) configured to communicate with other devices via a wired connection according to a suitable wired communication protocol. In particular examples, the network interface 112d includes the wired interface 112f and excludes the wireless interface 112e. In some examples, the electronic device 112 omits the network interface 112d entirely and transmits and receives media content and / or other data via another communication path (eg, input / output 111).

[0045] Audio component 112g is configured to process and / or filter data including media content received by electronic device 112 (e.g., via input / output 111 and / or network interface 112d) to generate an output audio signal. In some examples, audio processing component 112g comprises, for example, one or more digital-to-analog converters (DACs), audio pre-processing components, audio enhancement components, digital signal processors (DSPs), and / or other suitable audio processing components, modules, circuits, etc. In particular examples, one or more of audio processing components 112g may comprise one or more subcomponents of processor 112a. In some examples, electronic device 112 omits audio processing component 112g. In some aspects, for example, processor 112a executes instructions stored in memory 112b to perform audio processing operations to generate an output audio signal.

[0046] The amplifiers 112h are configured to receive and amplify audio output signals generated by the audio processing component 112g and / or the processor 112a. The amplifiers 112h may comprise electronic devices and / or components configured to amplify the audio signals to a level sufficient to drive one or more of the transducers 114. In some examples, for example, the amplifiers 112h include one or more switching or class-D power amplifiers. However, in other examples, the amplifiers include one or more other types of power amplifiers (e.g., linear gain power amplifiers, class-A amplifiers, class-B amplifiers, class-AB amplifiers, class-C amplifiers, class-D amplifiers, class-E amplifiers, class-F amplifiers, class-G amplifiers, and / or class-H amplifiers, and / or other suitable types of power amplifiers). In particular examples, the amplifiers 112h comprise a suitable combination of two or more of the aforementioned types of power amplifiers. Furthermore, in some examples, individual ones of the amplifiers 112h correspond to individual ones of the transducers 114. However, in other examples, the electronics 112 includes a single one of the amplifiers 112h configured to output an amplified audio signal to the plurality of transducers 114. In some other examples, the electronics 112 omits the amplifier 112h.

[0047] The transducer 114 (e.g., one or more speakers and / or speaker drivers) receives the amplified audio signal from the amplifier 112h and renders or outputs the amplified audio signal as sound (e.g., audible sound waves having a frequency between approximately 20 Hertz (Hz) and 20 Kilohertz (kHz)). In some examples, the transducer 114 may comprise a single transducer. However, in other examples, the transducer 114 comprises multiple audio transducers. In some examples, the transducer 114 comprises multiple types of transducers. For example, the transducer 114 may include one or more low-frequency transducers (e.g., subwoofers, woofers), a mid-frequency transducer (e.g., mid-range transducer, mid-woofer), and one or more high-frequency transducers (e.g., one or more tweeters). As used herein, "low frequency" can generally refer to audible frequencies below about 500 Hz, "mid-range frequency" can generally refer to audible frequencies between about 500 Hz and about 2 kHz, and "high frequency" can generally refer to audible frequencies above 2 kHz. However, in certain examples, one or more of the transducers 114 comprises a transducer that does not adhere to the aforementioned frequency ranges. For example, one of the transducers 114 may comprise a mid-woofer transducer configured to output sound at frequencies between about 200 Hz and about 5 kHz.

[0048] By way of example, SONOS, Inc. currently offers (or has offered) certain playback devices for sale, including, for example, “SONOS ONE,” “PLAY:1,” “PLAY:3,” “PLAY:5,” “PLAYBAR,” “PLAYBASE,” “CONNECT:AMP,” “CONNECT,” and “SUB.” Additionally or alternatively, other suitable playback devices may be used to implement the example playback devices disclosed herein. Furthermore, as will be appreciated by those skilled in the art, playback devices are not limited to the examples described herein or to the SONOS product offerings. In some examples, for example, one or more of the playback devices 110 comprise wired or wireless headphones (e.g., over-ear headphones, earbuds, in-ear earphones). In other examples, one or more of the playback devices 110 comprise a docking station and / or interface configured to interact with a docking station for a personal mobile media playback device. In certain examples, the playback device may be integrated with another device or component, such as a television, a lighting fixture, or some other device for indoor or outdoor use. In some examples, the playback device omits a user interface and / or one or more transducers. For example, FIG. 1D is a block diagram of a playback device 110 p with input / output 111 and electronics 112 without a user interface 113 or transducer 114 .

[0049] FIG. 1E is a block diagram of a combined playback device 110q comprising playback device 110i (e.g., a subwoofer) ( FIG. 1A ) and ultrasonically bonded playback device 110a ( FIG. 1C ). In the illustrated example, playback devices 110a and 110i are separate playback devices 110 housed in separate housings. However, in some examples, combined playback device 110q comprises a single housing housing both playback devices 110a and 110i. Combined playback device 110q can be configured to process and reproduce sound differently than uncoupled playback devices (e.g., playback device 110a of FIG. 1C ) and / or paired or combined playback devices (e.g., playback devices 110l and 110m of FIG. 1B ). In some examples, for example, playback device 110a is a full-range playback device configured to render low-, mid-, and high-frequency audio content, and playback device 110i is a subwoofer configured to render low-frequency audio content. In some aspects, playback device 110a, when coupled with a first playback device, is configured to render only the mid- and high-frequency components of a particular audio content, while playback device 110i renders the low-frequency components of the particular audio content. In some examples, combined playback device 110q includes additional playback devices and / or another combined playback device.

[0050] c. A suitable Network Microphone Device (NMD) FIG. 1F is a block diagram of NMD 120a (FIGS. 1A and 1B). NMD 120a includes one or more audio processing components 124 (hereinafter “audio components 124”) and several components described with respect to playback device 110a (FIG. 1C), including processor 112a, memory 112b, and microphone 115. NMD 120a optionally includes other components also included in playback device 110a (FIG. 1C), such as user interface 113 and / or transducer 114. In some examples, NMD 120a is configured as a media playback device (e.g., one or more of playback devices 110) and further includes, for example, one or more of audio component 112g (FIG. 1C), amplifier 114, and / or other playback device components. In particular examples, NMD 120a comprises an Internet of Things (IoT) device, such as a thermostat, an alarm panel, a fire and / or a smoke detector, etc. In some examples, NMD 120a includes microphone 115, audio processing 124, and only some of the components of electronics 112 described above with respect to FIG. 1B. In some aspects, for example, NMD 120a includes processor 112a and memory 112b (FIG. 1B) while omitting one or more other components of electronics 112. In some examples, NMD 120a includes additional components (e.g., one or more sensors, a camera, a thermometer, a barometer, a hygrometer).

[0051] In some examples, an NMD can be incorporated into a playback device. FIG. 1G is a block diagram of playback device 110r including NMD 120d. Playback device 110r can include many or all of the components of playback device 110a and can further include microphone 115 and audio processing 124 (FIG. 1F). Playback device 110r may have an integrated control device 130c. Control device 130c can include, for example, a user interface (e.g., user interface 113 of FIG. 1B) configured to receive user input (e.g., touch input, voice input) without a separate control device. However, in other examples, playback device 110r receives commands from another control device (e.g., control device 130a of FIG. 1B).

[0052] Referring again to FIG. 1F , the microphone 115 is configured to acquire, capture, and / or receive sound from the environment (e.g., environment 101 of FIG. 1A ) and / or room in which the NMD 120a is located. The received sound may include, for example, voice utterances, audio played by the NMD 120a and / or another playback device, background sounds, ambient sounds, etc. The microphone 115 converts the received sound into electrical signals to generate microphone data. The audio processing 124 receives and analyzes the microphone data to determine whether the microphone data contains voice input. The voice input may include, for example, a wake-up word followed by an utterance containing a user request. As will be appreciated by those skilled in the art, a wake-up word is a word or other audio cue signifying user voice input. For example, when querying an AMAZON® VAS, a user may utter the wake-up word “Alexa.” Other examples include “Okay, Google” to invoke a Google® VAS and “Hey, Siri” to invoke an Apple® VAS.

[0053] After detecting the hotword, voice processing 124 monitors microphone data for an accompanying user request in the voice input. The user request may include, for example, a command to control a third-party device such as a thermostat (e.g., a NEST® thermostat), a lighting device (e.g., a PHILIPS HUE® lighting device), or a media playback device (e.g., a Sonos® playback device). For example, a user may utter the hotword “Alexa” (e.g., environment 101 of FIG. 1A ) followed by the utterance “Set the thermostat to 68 degrees” to set the temperature in their home. A user may utter the same hotword followed by the utterance “Turn on the living room” to turn on a lighting device in the living room area of ​​their home. A user may similarly speak the hotword followed by a request to play a particular song, album, or music playlist on a playback device in their home.

[0054] d. Appropriate control devices FIG. 1H is a partial schematic diagram of control device 130a (FIGS. 1A and 1B). As used herein, the term "control device" can be used interchangeably with "controller" or "control system." Among other features, control device 130a is configured to receive user input associated with media playback system 100 and, in response, cause one or more devices within media playback system 100 to perform an action or actions corresponding to the user input. In the illustrated example, control device 130a comprises a smartphone (e.g., iPhone®, Android phone) having media playback system controller application software installed. In some examples, control device 130a comprises, for example, a tablet (e.g., iPad®), a computer (e.g., laptop computer, desktop computer), and / or another suitable device (e.g., television, automobile audio head unit, IoT device). In particular examples, control device 130a comprises a dedicated controller for media playback system 100. In other examples, as described above with respect to FIG. 1G, control device 130a is incorporated into another device in media playback system 100 (e.g., one or more of playback device 110, NMD 120, and / or other suitable devices configured to communicate over a network).

[0055] Control device 130a includes electronics 132, a user interface 133, one or more speakers 134, and one or more microphones 135. Electronics 132 includes one or more processors 132a (hereinafter referred to as “processor 132a”), memory 132b, software components 132c, and a network interface 132d. Processor 132a can be configured to perform functions related to facilitating user access, control, and configuration of media playback system 100. Memory 132b can include data storage into which one or more of the software components executable by processor 112a to perform these functions can be loaded. Software component 132c can include applications and / or other executable software configured to facilitate control of media playback system 100. Memory 112b can be configured to store, for example, software component 132c, media playback system controller application software, and / or other data related to media playback system 100 and users.

[0056] Network interface 132d is configured to facilitate network communication between control device 130a and one or more other devices in media playback system 100 and / or one or more remote devices. In some examples, network interface 132d is configured to operate according to one or more appropriate communications industry standards (e.g., infrared, wireless, wired standards including IEEE 802.3, wireless standards including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G, LTE). Network interface 132d can be configured to transmit and / or receive data to, for example, playback device 110, NMD 120, other control devices 130, one of computing devices 106 of FIG. 1B, one or more other devices comprising the media playback system, etc. Transmitted and / or received data can include, for example, playback device control commands, state variables, playback zones and / or zone group configurations. For example, based on user input received at user interface 133, network interface 132d can transmit playback device control commands (e.g., volume control, audio playback control, audio content selection) from control device 130 to one or more of playback devices 110. Network interface 132d can also transmit and / or receive configuration changes such as, for example, adding / removing one or more playback devices 110 to / from a zone, adding / removing one or more zones to / from a zone group, forming a combined or integrated player, separating one or more playback devices from a combined or integrated player, among others. A description of adding zones and groups can be found below with respect to Figures 1I-1M.

[0057] The user interface 133 is configured to receive user input and can facilitate control of the media playback system 100. The user interface 133 includes media content techniques 133a (e.g., album art, lyrics, video), playback status indicators 133b (e.g., elapsed time and / or time remaining indicators), media content information area 133c, playback control area 133d, and zone indicators 133e. The media content information area 133c can include a display of relevant information (e.g., title, artist, album, genre, release year) about the currently playing media content and / or media content in a queue or playlist. The playback control area 133d can include selectable (e.g., via touch input and / or via a cursor or another suitable selector) icons to cause one or more playback devices in a selected playback zone or zone group to perform playback actions, such as play or pause, fast forward, rewind, skip next, skip previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit crossfade mode, etc. Playback control area 133d may also include selectable icons for changing equalization settings, playback volume, and / or other appropriate playback operations. In the illustrated example, user interface 133 comprises a display presented on a touchscreen interface of a smartphone (e.g., iPhone®, Android phone). However, in some examples, user interfaces of various formats, styles, and interactive sequences may alternatively be implemented on one or more networked devices to provide equivalent control access to a media playback system.

[0058] One or more speakers 134 (e.g., one or more transducers) may be configured to output audio to a user of control device 130a. In some examples, one or more speakers comprise individual transducers configured to output corresponding low, mid, and / or high frequencies. In some aspects, for example, control device 130a is configured as a playback device (e.g., one of playback devices 110). Similarly, in some examples, control device 130a is configured as an NMD (e.g., one of NMDs 120) and receives voice commands and other audio via one or more microphones 135.

[0059] The one or more microphones 135 may comprise, for example, one or more condenser microphones, electret condenser microphones, dynamic microphones, and / or other suitable types of microphones or transducers. In some examples, two or more of the microphones 135 are positioned to capture location information of an audio source (e.g., voice, audible sound) and / or configured to facilitate filtering of background noise. Furthermore, in certain examples, the control device 130a is configured to operate as a playback device and an NMD. However, in other examples, the control device 130a omits one or more speakers 134 and / or one or more microphones 135. For example, the control device 130a may comprise a device (e.g., a thermostat, an IoT device, a network device) that includes a portion of the electronics 132 and a user interface 133 (e.g., a touchscreen) without a speaker or microphone.

[0060] Proper playback device configuration 1I-1M show exemplary configurations of playback devices in zones and zone groups. Referring first to FIG. 1M, in one example, a single playback device can belong to a zone. For example, playback device 110g in the second bedroom 101c (FIG. 1A) may belong to Zone C. In some implementations described below, multiple playback devices can be "combined" to form a "combined pair," which together form a single zone. For example, playback device 110l (e.g., the left playback device) can be combined with playback device 110j (e.g., the right playback device) to form Zone A. The combined playback devices may have different playback responsibilities (e.g., channel responsibilities). In another implementation described below, multiple playback devices can be merged to form a single zone. For example, playback device 110h (e.g., the front playback device) can be merged with playback device 110i (e.g., a subwoofer) and playback devices 110j and 110k (e.g., left and right surround speakers, respectively) to form a single Zone D. In another example, playback devices 110g and 110h can be merged to form merged group or zone group 108b. Merged playback devices 110g and 110h may not be specifically assigned different playback responsibilities. That is, merged playback devices 110h and 110i can play audio content the same way as if they were not merged, apart from playing audio content synchronously.

[0061] Each zone within media playback system 100 may be provided for control as a single user interface (UI) entity. For example, Zone A may be provided as a single entity called the master bathroom. Zone B may be provided as a single entity called the master bedroom. Zone C may be provided as a single entity called the second bedroom.

[0062] Combined playback devices may have different playback responsibilities, such as responsibility for specific audio channels. For example, as shown in FIG. 1-I, playback devices 110l and 110m may be combined to create or enhance a stereo effect for audio content. In this example, playback device 110l may be configured to play the left channel audio component, while playback device 110k may be configured to play the right channel audio component. In some implementations, such stereo combining may be referred to as "pairing."

[0063] Furthermore, combined playback devices may have additional and / or different respective speaker drivers. As shown in FIG. 1J, a playback device 110h labeled Front may be combined with a playback device 110i labeled SUB. The Front device 110h may be configured to render a mid- to high-frequency range, and the SUB device 110i may be configured to render low frequencies. However, when uncombined, the Front device 110h may be configured to render the entire frequency range. As another example, FIG. 1K shows the Front device 110h and the SUB device 110i further combined with the left playback device 110j and the right playback device 110k, respectively. In some implementations, the right device 110j and the left device 110k may be configured to form a surround or "satellite" channel in a home theater system. The combined playback devices 110h, 110i, 110j, and 110k may form a single zone D (FIG. 1M).

[0064] Merged playback devices may not have assigned playback responsibilities and may be capable of rendering the full range of audio content for which each playback device is capable. Nevertheless, merged devices may be represented as a single UI entity (i.e., a zone, as described above). For example, playback devices 110a and 110n in the master bathroom have a single UI entity for Zone A. In one example, playback devices 110a and 110n can each output the full range of audio content for which each playback device 110a and 110n is capable in sync.

[0065] In some examples, an NMD is combined or merged with another device to form a zone. For example, NMD 120b may be combined with playback device 110e, which together form zone F, referred to as the living room. In other examples, a standalone network microphone device may itself be within a zone. However, in other examples, a standalone network microphone device may not be associated with a zone. Further details regarding associating network microphone devices and playback devices as designated or default devices can be found, for example, in the above-referenced U.S. patent application Ser. No. 15 / 438,749.

[0066] Zones of individual devices, combined devices, and / or merged devices may be grouped to form zone groups. For example, referring to FIG. 1M, zone A may be grouped with zone B to form zone group 108a containing the two zones. Similarly, zone G may be grouped with zone H to form zone group 108b. As another example, zone A may be grouped with one or more other zones C1. Zones A-I may be grouped and ungrouped in numerous ways. For example, three, four, five, or more (e.g., all) of zones A-I may be grouped. Once grouped, zones of individual and / or combined playback devices may play audio in synchronization with one another, as described in the previously referenced U.S. Patent No. 8,234,395. Playback devices may also be dynamically grouped and ungrouped to form new or different groups that play audio content in synchronization.

[0067] In various implementations, a zone within an environment may be a combination of the default names of the zones within the group or the names of the zones within the zone group. For example, zone group 108b may be assigned a name such as "dining + kitchen," as shown in FIG. 1M. In some examples, a zone group may be given a unique name selected by the user.

[0068] Certain data may be stored in the memory of the playback device (e.g., memory 112b of FIG. 1C) as one or more state variables that are periodically updated and used to describe the state of the playback zone, the playback device, and / or the zone group associated with the playback zone. The memory may also contain data that is associated with the state of other devices in the media system and that is shared from time to time between devices so that one or more of the devices have the most current data associated with the system.

[0069] In some examples, the memory may store instances of various variable types associated with states. The variable instances may be stored with an identifier (e.g., a tag) corresponding to the type. For example, a particular identifier may be a first type "a1" to identify a playback device in a zone, a second type "b1" to identify playback devices that can be combined within the zone, and a third type "c1" to identify a zone group to which the zone can belong. As a related example, an identifier associated with the second bedroom 101c may indicate that the playback device is the only playback device in zone C and not within a zone group. An identifier associated with Den may indicate that Den is not grouped with other zones but includes combined playback devices 110h-110k. An identifier associated with the dining room may indicate that the dining room is part of the dining + kitchen zone group 108b and that devices 110b and 110d are grouped together (FIG. 1L). An identifier associated with the kitchen may indicate the same or similar information by virtue of the kitchen being part of the dining + kitchen zone group 108b. Other exemplary zone variables and identifiers are described below.

[0070] In yet another example, media playback system 100 may include variables or identifiers representing other associations of zones and zone groups, such as identifiers associated with areas, as shown in FIG. 1M. Areas can include clusters of zone groups and / or zones not within a zone group. For example, FIG. 1M shows upper area 109a including zones A-D and lower area 109b including zones E-I. In one aspect, an area can be used to refer to a zone group and / or cluster of zones that share one or more zones and / or zone groups of another cluster. In another aspect, this is distinct from a zone group that does not share a zone with another zone group. Further examples of techniques for implementing areas can be found, for example, in U.S. Patent Application No. 15 / 682,506, filed August 21, 2017, entitled "Name-Based Room Association," and U.S. Patent No. 8,483,853, filed September 11, 2007, entitled "Control and Manipulation of Groupings in a Multi-Zone Media System." Each of these applications is incorporated herein by reference in its entirety. In some examples, media playback system 100 may not implement areas, in which case the system may not store variables associated with areas.

[0071] III. Playback of Generated Media Content FIG. 2 is a functional block diagram of a system 200 for the playback of generated media content. As previously mentioned, generated media content can include any media content (e.g., audio, video, audiovisual output, tactile output, or any other media content) that is dynamically created, synthesized, and / or modified by a non-human, rule-based process, such as an algorithm or model. This creation or modification can occur for playback in real time or near real time. Additionally or alternatively, generated media content can be generated or modified asynchronously (e.g., before playback is requested), and specific items of generated media content can then be selected for playback at a later time. As used herein, a "generative media module" includes any system capable of generating generated media content based on one or more inputs, whether implemented with software, a physical model, or a combination thereof. In some examples, such generated media content includes new media content that can be created entirely new or by mixing, combining, manipulating, or otherwise modifying one or more existing pieces of media content. As used herein, a "generated media content model" includes any algorithm, schema, or set of rules that can be used to generate new generated media content using one or more inputs (e.g., sensor data, artist-provided parameters, media segments such as audio clips or samples, etc.). In examples, a generative media module can use a variety of different generated media content models to generate different generated media content. In some cases, an artist or other collaborator can interact with, create, and / or update a generated media content model to generate particular generated media content.Although some examples throughout this description refer to audio content, the principles disclosed herein may, in some examples, be applied to other types of media content, such as video, audiovisual, tactile, or other.

[0072] 2, system 200 includes a produced media group coordinator 210 that communicates with produced media group members 250a and 250b, as well as a sensor data source 218, a media content source 220, and a control device 130. Such communication may be performed over network 102, which, as previously described, may include any suitable wired or wireless network connection or combination thereof (e.g., a WiFi network, Bluetooth, a Z-Wave network, ZigBee, an Ethernet connection, a Universal Serial Bus (USB) connection, etc.).

[0073] One or more remote computing devices 106 can also communicate with the group coordinator 210 and / or group members 250a and 250b via the network 102. In various examples, the remote computing device 106 can be a cloud-based server associated with a device manufacturer, a media content provider, a voice assistant service, or other suitable entity. As shown in FIG. 2, the remote computing device 106 can include a generated media module 214. As described in more detail elsewhere herein, the remote computing device 106 can generate generated media content remotely from local devices (e.g., the coordinator 210 and members 250a and 250b). The generated media content can then be transmitted to one or more local devices for playback. Additionally or alternatively, the generated media content can be generated in whole or in part via local devices (e.g., the group coordinator 210 and / or group members 250a and 250b). In some examples, the group coordinator 210 may itself be a remote computing device, communicatively coupled to group members 250a and 250b via a wide area network, and the devices need not be co-located within the same environment (e.g., home, business, etc.).

[0074] a. Example of generated media group behavior In the depicted example, the produced media group includes a produced media group coordinator 210 (also referred to herein as “coordinator device 210”) and first and second produced media group members 250 a and 250 b (also referred to herein collectively as “first member device 250 a,” “second member device 250 b,” and “member devices 250”). Optionally, one or more remote computing devices 106 may also form part of the produced media group. In operation, these devices may communicate with each other and / or other components (e.g., sensor data source 218, control device 130, media content source 220, or any other suitable data source or component) to facilitate the generation and playback of produced media content.

[0075] In various examples, some or all of devices 210 and / or 250 may be co-located in the same environment (e.g., in the same home, in a store, etc.) In some examples, at least some of devices 210 and / or 250 may be remote from one another, e.g., in different homes, in different cities, etc.

[0076] 1A-1H, the coordinator device 210 and / or the member device 250 may include some or all of the components of the playback device 110 or the network microphone device 120. For example, the coordinator device 210 and / or the member device 250 may optionally include playback components 212 (e.g., transducers, amplifiers, audio processing components, etc.), or such components may be omitted in some cases.

[0077] In some examples, the coordinator device 210 is a playback device itself and therefore can also operate as a member device 250. In other examples, the coordinator device 210 can be connected to one or more member devices 250 (e.g., via a direct wired connection or via network 102), but the coordinator device 210 does not itself play the produced media content. In various examples, the coordinator device 210 can be implemented on a bridge-like device on a local network, a playback device that is not itself part of a produced media group (i.e., the playback device does not itself play the produced media content), and / or a remote computing device (e.g., a cloud server).

[0078] In various examples, one or more of the devices may include a generated media module 214 thereon. Such generated media module 214 may generate new composite media content based on one or more inputs, for example, using an appropriate generated media content model. As shown in FIG. 2 , in some examples, the coordinator device 210 may include a generated media module 214 for generating generated media content, which may then be transmitted to member devices 250 a and 250 b for simultaneous and / or synchronized playback. Additionally or alternatively, some or all of the member devices 250 (e.g., member device 250 b shown in FIG. 2 ) may include a generated media module 214, which may be used by member device 250 to locally generate generated media content based on one or more inputs. In various examples, the generated media content may be generated via the remote computing device 106, optionally using one or more input parameters received from a local device. This generated media content may then be transmitted to one or more local devices for coordination and / or playback.

[0079] In some examples, at least some of the member devices 250 do not include a generated media module 214 therein. Alternatively, in some cases, each member device 250 can include a generated media module 214 therein and can be configured to generate generated media content locally. In at least some examples, none of the member devices 250 include a generated media module 214 therein. In such cases, generated media content can be generated by the coordinator device 210. Such generated media content can then be transmitted to the member devices 250 for simultaneous and / or synchronized playback.

[0080] 2, coordinator device 210 further includes coordination component 216. As described in more detail herein, in some cases, coordinator device 210 can facilitate playback of produced media content via multiple different playback devices (which may or may not include coordinator device 210 itself). In operation, coordination component 216 is configured to facilitate synchronization of both produced media creation (e.g., using one or more produced media modules 214, which may be distributed among various devices) and produced media playback. For example, coordinator device 210 can transmit timing data to member devices 250 to facilitate synchronized playback. Additionally or alternatively, the coordinator device 210 may send input, generated media model parameters, or other data regarding the generated media modules 214 to one or more member devices 250 so that the member devices 250 can generate the generated media locally (e.g., using locally stored generated media modules 214) and / or so that the member devices 250 can update or modify the generated media modules 214 based on input received from the coordinator device 210.

[0081] As described in more detail elsewhere herein, the generated media module 214 can be configured to generate generated media based on one or more inputs using a generated media content model. The inputs can include sensor data (e.g., as provided by a sensor data source 218), user input (e.g., as received from the control device 130 or via direct user interaction with the coordinator device 210 or a member device 250), and / or a media content source 220. For example, the generated media module 214 can generate and continuously modify the generated audio by adjusting various characteristics of the generated audio based on one or more input parameters (e.g., sensor data related to one or more users of the devices 210, 250).

[0082] b. Exemplary Media Content Sources Media content source 220, in various examples, can include one or more local and / or remote media content sources. For example, media content source 220 can include one or more local audio sources 105 (e.g., audio received via an input / output connection from a mobile device (e.g., a smartphone, tablet, laptop computer) or another suitable audio component (e.g., a television, desktop computer, amplifier, phonograph, Blu-ray player, memory storing digital media files), etc.) as described above. Additionally or alternatively, media content source 220 can include one or more remote computing devices accessible via a network interface (e.g., via communication over network 102). Such remote computing devices can include, for example, individual computers or servers, such as media streaming service servers, that store audio and / or other media content, etc.

[0083] In various examples, media available via media content source 220 may include pre-recorded audio segments in the form of complete sounds, songs, portions of songs (e.g., samples), or any audio component (e.g., pre-recorded audio of a particular instrument, synthesized beats or other audio segments, non-musical audio such as spoken word or natural sounds, etc.). In operation, such media may be utilized by generated media module 214 to generate generated media content, for example, by combining, mixing, overlaying, manipulating, or otherwise modifying the retrieved media content to generate new generated media content for playback via one or more devices. In some examples, the generated media content may take the form of a combination of pre-recorded audio segments (e.g., pre-recorded songs, spoken word recordings, etc.) and new synthesized audio that is created and overlaid with the pre-recorded audio. As used herein, “generated media content” or “generated media content” may include any such combination.

[0084] c. Exemplary Generative Media Module As previously mentioned, the generative media module 214 may include any system capable of generating generative media content based on one or more inputs, whether instantiated in software, a physical model, or a combination thereof. In various examples, the generative media module 214 may utilize a generative media content model, which may include one or more algorithms or mathematical models that determine how media content is generated based on associated input parameters. In some cases, the algorithms and / or mathematical models themselves may be updated over time, for example, based on instructions received from one or more remote computing devices (e.g., a cloud server associated with a music service or other entity), or based on input received from other group member devices in the same or different environments, or any other suitable input. In some examples, various devices in a group may have different generative media modules 214 thereon, for example, with a first member device having a different generative media module 214 than a second member device. In other cases, each device in a group having a generative media module 214 may include substantially the same model or algorithm.

[0085] Any suitable algorithm or combination of algorithms may be used to generate the generative media content. Examples of such algorithms include those using machine learning techniques (e.g., generative adversarial networks, neural networks, etc.), formal grammars, Markov models, finite-state automata, and / or any algorithms implemented in currently available offerings such as JukeBox by OpenAI, AWS DeepComposer by Amazon, Magenta by Google, and Amper AI by Amper Music. In various examples, the generative media module 214 may utilize any suitable generative algorithm currently existing or developed in the future.

[0086] In line with the above description, generating generative media content (e.g., audio content) can include modifying various characteristics of the media content in real time and / or algorithmically generating new media content in real time or near real time. In the context of audio content, this can be accomplished by storing several audio samples in a database (e.g., within the media content source 220), which can be remotely located and accessible by the coordinator device 210 and / or the member devices 250 via the network 102, or the audio samples can be maintained locally on the devices 210, 250 themselves. The audio samples can be associated with one or more metadata tags corresponding to one or more audio characteristics of the sample. For example, a given sample can be associated with metadata tags indicating that the sample contains audio of a particular frequency or frequency range (e.g., bass / midrange / treble), or a particular instrument, genre, tempo, key, release date, geographic region, timbre, reverb, distortion, sonic texture, or any other audio characteristic that becomes apparent.

[0087] In operation, the generative media module 214 (e.g., of the coordinator device 210 and / or the second member device 250b) can search for specific audio samples based on their associated tags and blend the audio samples to create the generated audio. The generated audio can evolve in real time as the generative media module 214 searches for audio samples with different tags and / or different audio samples with the same or similar tags. The audio samples that the generative media module(s) 214 search for can depend on one or more inputs, such as sensor data, time of day, geographic location, weather, or various user inputs, such as mood selection, or physiological inputs, such as heart rate. In this manner, as the inputs change, the generated audio also changes. For example, if a user selects a calming or relaxing mood input, the generative media module(s) 214 can search for and blend audio samples with tags corresponding to audio content that the user finds calming or relaxing. Examples of such audio samples can include audio samples tagged as low tempo or low harmonic complexity, or audio samples predetermined and tagged as calm and relaxing. In some examples, audio samples may be identified as calming or relaxing based on an automated process that analyzes the temporal and spectral content of the signal. Other examples are possible as well. In any of the examples herein, the generated media module 214 may adjust the characteristics of the generated audio by searching for and mixing audio samples associated with different metadata tags or other suitable identifiers.

[0088] Modifying the characteristics of the generated audio can include manipulating one or more of the volume, balance, removal of particular instruments or tones, changing the tempo, gain, reverb, spectral equalization, timbre, or sound texture of the audio, etc. In some examples, the generated audio can be played back differently on different devices, such as by emphasizing particular characteristics of the generated audio on a particular playback device closest to the user. For example, the closest playback device can emphasize a particular instrument, beat, tone, or other characteristic, while the remaining playback devices can serve as background audio sources.

[0089] As described elsewhere herein, the media content module 214 can be configured to generate media intended to steer the user's mood and / or physiological state in a desired direction. In some examples, the user's current state (e.g., mood, emotional state, activity level, etc.) is constantly and / or repeatedly monitored or measured (e.g., at predetermined intervals) to ensure that the user's current state is moving toward the desired state or at least not moving in the opposite direction to the desired state. In such examples, the generated audio content can be altered to move the user's current state toward a desired end state.

[0090] In any of the examples herein, the generative media module may use hysteresis to avoid rapid adjustments to the generated audio that could adversely affect the listening experience. For example, if the generative media module changes media based on user position input relative to the playback device, the playback device may rapidly change the generated audio in any of the methods described herein when the user moves closer to or further away from the playback device. Such abrupt adjustments may be unpleasant for the user. To reduce these rapid adjustments, the generative media module 214 may be configured to use hysteresis by delaying adjustments to the generated audio for a predetermined period of time when user movement or other activity triggers an adjustment. For example, if the playback device detects that the user has moved within a threshold distance of the playback device, instead of immediately performing one of the adjustments described above, the playback device may wait a predetermined amount of time (e.g., several seconds) before making the adjustment. If the user remains within the threshold distance after the predetermined amount of time, the playback device may proceed to adjust the generated audio. However, if the user does not remain within the threshold distance after the predetermined amount of time, the generative media module 214 may refrain from adjusting the generated audio. The generated media module 214 can similarly apply hysteresis to other generated media adjustments described herein.

[0091] FIG. 3 shows a flowchart of a process 300 for generating generated audio content using various input parameters. In various examples, one or more of these input parameters can be modified based on user input. For example, an artist can select various parameters, constraints, or available audio segments shown in FIG. 2, and these selections can at least partially determine the final output of the generated audio content. As previously described, such generated media modules may be stored and manipulated on one or more playback devices for local playback (e.g., via the same playback device and / or via other playback devices communicatively coupled via a local area network). Additionally or alternatively, such generated media modules may be stored and manipulated on one or more remote computing devices, with the resulting output transmitted to one or more remote devices for playback over a wide area network.

[0092] As shown, the process begins at block 302 and proceeds to a clock / metronome at block 304, where inputs of a tempo 306 and a time signature 308 are received. The tempo 306 and time signature 308 can be selected by the artist or can be automatically determined or generated using a model. The process proceeds to block 310, where a chord change can be triggered, receiving as input a chord change frequency parameter 312. The artist may choose to have a higher chord change frequency in music intended for a higher energy experience (e.g., dance music, uplifting atmospheric music, etc.). Conversely, a lower chord change frequency may be associated with a lower energy output (e.g., quieter music).

[0093] In block 314, a chord is selected from an available chord segment 316. A number of chord information parameters 318, 320, 322 may also be provided as inputs to the chord segment 316. These inputs may be used to determine the particular chord to be played next and output as block 324. In some instances, the artist may provide information for each chord, such as a weighting, how often that particular chord should be used, etc.

[0094] Next, in block 326, a chord variation is selected based at least in part on the harmonic complexity parameter that serves as an input. The harmonic complexity parameter 328 may be adjusted or selected by the artist, or may be automatically determined. In general, a higher harmonic complexity parameter may be associated with a higher energy audio output, and a lower harmonic complexity parameter may be associated with a lower energy audio output. In some cases, the harmonic complexity parameter may include inputs such as chord inversions, voicing, and harmonic density.

[0095] In block 330, the process gets the root of the chord and selects bass segments to play from available bass segments 334 in block 332. These bass segments then undergo bass processing 336, where equalization, filtering, timing, and other processing may be performed.

[0096] Returning to the chord variation of block 326, the process continues separately to block 338 to play a selected harmony from among the available harmony segments 340. This harmony segment then undergoes bass processing 342. Similar to low bass processing, the harmony segment bass processing 342 can include equalization, filtering, timing, etc.

[0097] Returning to the selected chord 324, the process continues to filter the melody notes individually at block 344 utilizing the input of melodic constraints 346. The output at block 348 is the melody notes available for playback. The melodic constraints 346 may be provided by the artist and may, for example, specify which notes to play or not, limit the melodic range, or provide other such constraints that may depend on the particular selected chord 324.

[0098] In block 350, the process determines which melody note (among available melody notes 348) to play. This determination can be made automatically based on model values, artist-provided input, randomization effects, or any other suitable input. In the illustrated example, one input is from trigger melody note block 352, which is based on a melodic density parameter 354. The artist can provide the melodic density parameter 354, which in part determines how complex and / or high-energy the audio output will be. Based on that parameter, melody notes may be triggered more or less frequently and at specific times using block 352, which is input to block 350, to determine which melody note to play. In various examples, the output of block 350 may be provided as an input to block 350 in the form of a feedback loop, such that the next melody note selected in block 350 depends at least in part on the melody note last selected in block 350. Next, in block 356, a melody segment is selected from available melody segments 358, after which the melody segment undergoes bass processing 360.

[0099] Returning to the start at block 302, the process separately proceeds to block 362 to play non-musical content. This may be, for example, nature sounds, speech sounds, or other such non-musical content. Various non-musical segments 364 may be stored and played. These non-musical content segments may also undergo bass processing at block 366.

[0100] The outputs of these various paths (e.g., selected bass segment(s), harmony segment(s), melody segment(s), and / or non-musical segment(s)) may each undergo separate bus processing before being combined at block 368 via a mixing and mastering process, where combined levels may be set, various filters may be applied, relative timing may be established, and any other appropriate processing steps may be performed before the generated audio content is output at block 370. In various examples, some of the paths may be omitted entirely. For example, the generative media module may omit the option to play non-musical content along with the generated musical content. The process 300 shown in FIG. 3 is exemplary only, and those skilled in the art will recognize that appropriate modifications can be made to the process 300 shown therein, and further, that there are numerous suitable alternative processes that may be used to generate generated media content.

[0101] 4 is an exemplary architecture for storing and retrieving generated media content. In this example, the generated media content includes a variety of individual tracks (each having multiple variations related to energy level or another parameter) that can be selected and played in various orders and groupings depending on certain input parameters.

[0102] As shown, generated media content 404 may be stored as one or more audio files associated with global generated media content metadata 402. Such metadata may include, for example, a global tempo (e.g., beats per minute), a global trigger frequency (e.g., how often to check for changes in an input parameter), and / or a global crossfade duration (e.g., the time to fade between different selected energies).

[0103] Within the generated media content 404 are multiple different tracks 406, 408, and 410. During operation, these tracks can be selected and played in various arrangements (e.g., randomized grouping with some kind of overlay, or played according to a predetermined sequence, etc.). In some examples, the generated media content 404, including tracks 406, 408, and 410, can be stored locally via one or more playback devices, while one or more remote computing devices can periodically transmit updated versions of the tracks, generated media content, and / or global generated media content metadata. In some examples, the remote computing devices can be periodically polled or queried by the playback devices, and in response to the queries or polls, the remote computing devices can provide updates to the generated media modules stored on the local playback devices.

[0104] For each track, there may be a corresponding subset of that track corresponding to a different energy level. For example, a first energy level (EL) for track 1 at 412, a second energy level for track 1 at 414, and an nth energy level for track 1 at 416. Each of these may include both metadata (e.g., metadata 418, 420, 422) and specific media files (e.g., media files 424, 426, 428) corresponding to the particular energy level. In some examples, each track may include multiple media files (e.g., media file 424) arranged in a particular manner, the arrangement and combination of which may be desired by the corresponding metadata (e.g., metadata 418). The media files may be in any suitable format, for example, playable via a playback device and / or streamable to a playback device for playback. In some examples, one or more of the media files 424, 426, 428 may be the output of the generative model shown in FIG. 3. The metadata may include, for example, tempo (if different from the global tempo), trigger frequency (if different from the global trigger frequency), sequence information (e.g., whether to play specific files sequentially, randomly, or with percentage weighting), crossfade duration (if different from the global crossfade), spatial information (e.g., to render audio content in space using multiple transducers), polynomial information (e.g., allowing multiple audio files to be played in this segment at once), and / or level (e.g., level adjustment in dB, or random within a predetermined range).

[0105] During operation, a target energy level can be determined using one or more input parameters (e.g., the number of people present in the room, the time of day, etc.). This determination can be made using the playback device and / or one or more remote computing devices. Based on this determination, specific media files corresponding to the determined energy level can be selected. The generative media module can then arrange and play those selected tracks according to the generative content model. This can include playing the selected tracks in a specific, predetermined order, playing them in a random or pseudo-random order, or any other suitable approach. The tracks can be played in an at least partially overlapping manner in some examples. It can be useful to vary the amount of overlap between tracks so that a casual listener does not hear a repeated loop of audio content, but instead perceives the generated audio as an endless stream of non-repeating audio.

[0106] While the example shown in FIG. 4 utilizes energy level as a parameter for distinguishing different generated audio content, in various examples, the particular variation or permutation of the generated audio content may vary along other dimensions (e.g., genre, time of day, associated user task, etc.).

[0107] d. Exemplary Sensor Data Sources and Other Input Parameters As previously discussed, the generated media module 214 may generate generated media based at least in part on input parameters, which may include sensor data (e.g., if received from a sensor data source 214) and / or other suitable input parameters. With respect to sensor input parameters, the sensor data source 214 may include data from any suitable sensor, wherever located relative to the generated media group and whatever values ​​measured thereby. Examples of suitable sensor data include physiological sensor data, such as data obtained from biometric sensors, wearable sensors, etc. Such data may include physiological parameters such as heart rate, respiration rate, blood pressure, brain waves, activity level, movement, body temperature, etc.

[0108] Suitable sensors include wearable sensors configured to be worn or carried by a user, such as a headset, a watch, a mobile device, a brain-machine interface (e.g., Neuralink), headphones, a microphone, or other similar devices. In some examples, the sensors may be non-wearable sensors or may be affixed to a fixed structure. The sensors may provide sensor data that may include, for example, data corresponding to brain activity, speech, position, movement, heart rate, pulse, body temperature, and / or sweating. In some examples, the sensors may correspond to multiple sensors. For example, as described elsewhere herein, the sensors may correspond to a first sensor worn by a first user, a second sensor worn by a second user, and a third sensor not worn by a user (e.g., affixed to a stationary body or structure). In such examples, the sensor data may correspond to multiple signals received from each of the first, second, and third sensors.

[0109] The sensors can be configured to acquire or generate information generally corresponding to the user's mood or emotional state. In one example, the sensor is a wearable brain-sensing headband, one of many examples of sensors described herein. Such a headband can include, for example, an electroencephalography (EEG) headband having multiple sensors thereon. In some examples, the headband can correspond to any of the Muse™ headbands (InteraXon; Toronto, Canada). The sensors can be positioned at various locations around the inner surface of the headband to correspond to, for example, different brain anatomies of the user (e.g., the frontal, parietal, temporal, and sphenoid bones). In this manner, each sensor can receive different data from the user. Each of the sensors can correspond to an individual channel that can be streamed from the headband to system device 210 and / or 250. Such sensor data can be used to detect the user's mood, for example, by classifying the frequency and intensity of various brain waves or by performing other analyses. Further details on using a brain-sensing headband for generating audio content can be found in commonly owned U.S. Patent Application No. 62 / 706,544, filed August 24, 2020, entitled MOOD DETECTION AND / OR INFLUENCE VIA AUDIO PLAYBACK DEVICES, which is incorporated herein by reference in its entirety.

[0110] In some examples, sensor data sources 218 include data obtained from networked device sensor data (e.g., Internet of Things (IoT) sensors, such as networked lights, cameras, temperature sensors, thermostats, presence detectors, microphones, etc.) Additionally or alternatively, sensor data sources 218 may include environmental sensors (e.g., measuring or displaying weather, temperature, time / day / week / month, etc.).

[0111] In some examples, the generated media module 214 can utilize inputs in the form of playback device capabilities (e.g., number and type of transducers, output power, other system architecture), device location (e.g., location relative to one or more users, relative to other playback devices). Further examples of generating and modifying generated audio as a result of user and device location are described in more detail in commonly owned U.S. patent application Ser. No. 62 / 956,771, filed January 3, 2020, and entitled "GENERATIVE MUSIC BASED ON USER LOCATION," which is incorporated herein by reference in its entirety. Additional inputs can include device states of one or more devices in a group, such as thermal state (e.g., if a particular device is in danger of overheating, the generated content can be modified to lower the temperature), battery level (e.g., bass output can be reduced in a portable playback device with a low battery level), and coupling state (e.g., whether a particular playback device is configured as part of a stereo pair, coupled with a sub, configured as part of a home theater setup, etc.). Any other suitable device characteristics or states may similarly be used as inputs for the generation of generated media content.

[0112] Another exemplary input parameter includes user presence; for example, when a new user enters a space playing generated audio, the user's presence can be detected (e.g., via a proximity sensor, beacon, etc.) and the generated audio can be modified in response. The modification can be based on the number of users (e.g., presenting ambient meditative audio for one user, relaxing music for two to four users, and party or dance music for more than four users). The modification can also be based on the identities of users present (e.g., user profiles based on user characteristics, listening history, or other such indicia).

[0113] In one example, a user may wear a biometric device that can measure various biometric parameters, such as the user's heart rate or blood pressure, and report those parameters to devices 210 and / or 250. The generated media modules 214 of these devices 210 and / or 250 may use these parameters to further adapt the generated audio, such as by increasing the tempo of the music in response to detecting a high heart rate (which may indicate that the user is engaged in high athletic activity) or by decreasing the tempo of the music in response to detecting high blood pressure (which may indicate that the user is stressed and could benefit from calming music).

[0114] In yet another example, one or more microphones of a playback device (e.g., microphone 115 of FIG. 1F) can detect the user's voice. The captured voice data can then be processed to, for example, determine the user's mood, age, or gender, identify a particular user among multiple users in a home, or any other such input parameter. Other examples are possible as well.

[0115] e. Examples of coordination between group members FIG. 5 is a functional block diagram illustrating data exchange in a system for playback of generated media content. For illustrative purposes, the system 500 shown in FIG. 5 includes interactions between a coordinator device 210 and a member device 250b. However, the interactions and processes described herein can be applied to interactions involving multiple additional coordinator devices 210 and / or member devices 250b. As shown in FIG. 5, the coordinator device 210 includes a generated media module 214a that receives inputs including input parameters 502 (e.g., sensor data, media content, model parameters of the generated media module 214a, or other such inputs) and clock and / or timing data 504. In various examples, the clock and / or timing data 504 can include synchronization signals for synchronizing playback and / or synchronizing generated media being generated by various devices in a group. In some examples, the clock and / or timing data 504 can be provided by an internal clock, processor, or other such component housed within the coordinator device 210 itself. In some examples, the clock and / or timing data 504 may be received from a remote computing device via a network interface.

[0116] Based on these inputs, the generative media module 214a may output generated media content 404a. Optionally, the output generated media content 404a may itself serve as input to the generative media module 214a in the form of a feedback loop. For example, the generative media module 214a may generate subsequent content (e.g., audio frames) using a model or algorithm that depends at least in part on previously generated content.

[0117] In the illustrated example, member device 250b similarly includes generated media module 214b, which may be substantially identical to generated media module 214a of coordinator device 210, or may differ in one or more aspects. Generated media module 214b may similarly receive input parameters 502 and clock and / or timing data 504. These inputs may be received from coordinator device 210, from other member devices, from other devices on the local network (e.g., a locally networked smart thermostat providing temperature data), and / or from one or more remote computing devices (e.g., a cloud server providing clock and / or timing data 504, or weather data, or any other such input). Based on these inputs, generated media module 214b may output generated media content 404b. This generated generated media content 404b may, optionally, be fed back to generated media module 214b as part of a feedback loop. In some examples, the generated media content 404b can include or consist of the generated media content 404a (generated via the coordinator device 210) that is sent over the network to the member device 250b. In other cases, the generated media content 404b can be generated independently and separately from the generated media content 404a generated via the coordinator device 210.

[0118] The generated media content 404a and 404b can then be played back via devices 210 and 250b themselves and / or for playback by other devices in the group. In various examples, the generated media content 404a and 404b can be configured for simultaneous and / or synchronous playback. In some cases, the generated media content 404a and 404b may be substantially identical or similar to one another, with each generated media module 214 utilizing the same or similar algorithms and the same or similar inputs. In other examples, the generated media content 404a and 404b may be different from one another, while still being configured for synchronized or simultaneous playback.

[0119] f. Exemplary Generative Media Using a Distributed Architecture As previously mentioned, generating media content can be computationally intensive and, in some cases, may be impractical to perform entirely on the local playback device alone. In some examples, a generated media module of the local playback device can request generated media content from generated media modules stored on one or more remote computing devices (e.g., cloud servers). The request can include or be based on specific input parameters (e.g., sensor data, user input, contextual information, etc.). In response to the request, the remote generated media module can stream the specific generated media content to the local device for playback. The specific generated media content provided to the local playback device can change over time depending on the specific input parameters, the configuration of the generated media module, or other such parameters. Additionally or alternatively, the playback device can store individual tracks for playback (e.g., different variations in the track are associated with different energy levels, as shown in FIG. 4). The remote computing device can then periodically provide new files for the updated tracks to the local playback device for playback or provide updates to the generated media module that determines when and how to play specific files stored locally on the playback device.

[0120] In this manner, the tasks required to generate and play back generated audio are distributed between one or more remote computing device(s) and one or more local playback devices. By performing at least some of the computationally intensive tasks associated with generating new media content on the remote computing devices and, optionally, by reducing the need for real-time computation, overall efficiency can be improved. By generating a discrete number of alternative tracks or track variations according to a particular media content model prior to playback via the remote computing devices, the local playback device can request and receive particular variations based on real-time or near-real-time input parameters (e.g., sensor data). For example, the remote computing devices can generate different versions of media content, and the playback device can request particular versions in real time based on the input parameters. This results in playback of appropriate generated media content based on real-time or near-real-time input parameters (e.g., sensor data) without requiring de novo generation of such media content performed in real time.

[0121] 6 is a schematic diagram of an exemplary distributed generative media playback system 600. As shown, an artist 602 can provide multiple media segments 604 and one or more generative content models 606 to a stored generative media module 214 via one or more remote computing devices. A media segment can correspond, for example, to a particular audio segment or seed (e.g., individual notes or chords, a short n-bar track, non-musical content, etc.). In some examples, the generative content model 606 can also be provided by the artist 602. This can include providing the entire model, or the artist 602 can provide input to the model 606 by, for example, modifying or adjusting certain aspects (e.g., tempo, melodic constraints, harmonic complexity parameters, chord change density parameters, etc.).

[0122] The generative media module 214 may receive both a media segment 604 and one or more input parameters 502 (as described elsewhere herein). Based on these inputs, the generative media module 214 may output generated media. As shown in FIG. 6, the artist 602 may optionally audition the generative media module 214, for example, by receiving exemplary outputs based on inputs (e.g., media segments 604 and / or generative content models 606) provided by the artist 602. In some cases, the audition may play variations of the generated media content to the artist 602 in response to a variety of different input parameters (e.g., one version corresponding to a high energy level intended to create an exciting or uplifting effect, another version corresponding to a low energy level intended to create a calming effect, etc.). Based on the output from this auditioning step, the artist 602 may dynamically update settings of the media segment 604 and / or generative content model 606 until a desired output is achieved.

[0123] In the illustrated example, there may be an iteration in block 608 every n hours (or minutes, days, etc.), during which the generated media module 214 may generate multiple different versions of the generated media content. In the illustrated example, there are three versions: version A in block 610, version B in block 612, and version C in block 614. These outputs are stored (e.g., via a remote computing device) as generated media content 616. A particular version (in this example, version C as block 618) may be sent (e.g., streamed) to the local playback device 250 for playback. In some examples, the particular version may correspond to tracks 406, 408, and 410 shown in FIG. 4.

[0124] Although three versions are shown here by way of example, in practice there may be many more versions of generated media content generated via a remote computing device. The versions may vary along several different dimensions, such as being suitable for different energy levels, for different intended tasks or activities (e.g., studying versus dancing), for different times of day, or any other suitable variation.

[0125] In the illustrated example, playback device 250 can periodically request a particular version of the generated media content from a remote computing device. Such a request can be based, for example, on user input (e.g., user selection via a controller device), sensor data (e.g., the number of people present in a room, background noise level, etc.), or other suitable input parameters. As illustrated, input parameters 502 can optionally be provided to (or detected by) playback device 250. Additionally or alternatively, input parameters 502 can be provided to (or detected by) remote computing device 106. In some examples, playback device 250 transmits the input parameters to remote computing device 106, which provides the appropriate version to playback device 250 without playback device 250 specifically requesting a particular version.

[0126] g. Examples of Multi-Channel Media Content Creation and Playback In some examples, generated media content can take the form of multi-channel content. The channels can correspond to traditional audio distribution (e.g., left, right, ambient, height) or other distributions (e.g., a first channel of nature sounds and a second channel of rhythmic beats). Furthermore, generated media content can be included as one channel of multi-channel audio that also includes non-generated media content. In some cases, multi-channel playback of generated media content can present particular challenges, particularly with regard to synchronizing the playback of various channels across different playback devices in an environment. For example, the particular distribution of generated media content among different playback devices may change in real time based on particular inputs (e.g., sensor data, user input, or other contextual information), and synchronizing such playback given dynamic adjustments that rely solely on remote computing devices to determine playback responsibility can introduce undesirable latency.

[0127] Some examples of the present technology address these and other issues by providing all channels of multi-channel generated media content (e.g., multi-channel content including at least some generated media content) to each of multiple playback devices in an environment. FIG. 7 shows an exemplary distributed generated media playback system 700. System 700 may be similar to system 600 described above with respect to FIG. 6, and various implementations may include certain components omitted from FIG. 7. Some aspects of generated media content generation have been omitted; only generated media module 214 and the resulting generated media content 616 are shown here. However, in various implementations, any of the techniques or technologies described elsewhere herein or otherwise known to those skilled in the art may be incorporated into the generation of generated media content 616.

[0128] In various examples, the generated media content 616 may include multi-channel media content. The generated media content 616 may then be transmitted to the group coordinator 210, which may be a playback device or any other suitable device in the local environment, as previously described. The group coordinator 210 may communicate the generated media content 616 to each of a number of member devices or playback devices 250a, 250b, and 250c (collectively, “member devices” 250 or “playback devices 250”). Additionally, each playback device 250 may be configured to receive one or more input parameters 502. As previously described, the input parameters 502 may include any suitable input, such as user input (e.g., user selection via a controller device), sensor data (e.g., number of people present in a room, background noise level, time of day, weather data, etc.), or other suitable input parameters. In various examples, the input parameters 502 may optionally be provided to the playback device 250 and / or detected or determined by the playback device 250 itself.

[0129] In some examples, each channel of the multi-channel media content 616 is transmitted to both the group coordinator 210 and each playback device 250. The transmitted content may be divided into frames by the coordinator 210 before transmission, or may be transmitted in unencoded form (e.g., as a PCM signal). If the content 616 is encoded, it may be decoded at each playback device 250. Although each of the playback devices 250 receives each channel of the multi-channel media content, the playback devices 250 may have different playback responsibilities. For example, a first playback device 250a may be assigned to play only a first subset of the channels of the multi-channel media content, a second playback device 250b may be assigned to play a second subset of the channels, and a third playback device 250c may be assigned to play yet a third subset of the channels. These subsets may be completely different or may at least partially overlap. Furthermore, in addition to playing specific subsets of channels, various playback levels may also differ among the playback devices 250. For example, to create the effect of a rainstorm concentrated in one corner of a room, a channel of audio corresponding to the sound of rain may be played at a first level by a playback device directly in that corner, while a second playback device spaced away from the corner may play the rain sound channel at a lower level. In at least some examples, fewer than all channels of the multi-channel media content 616 are transmitted to each of the playback devices.

[0130] In some examples, to determine specific playback responsibilities and coordinate synchronized playback among various devices, the group coordinator 210 can send timing and / or playback responsibility information to the playback devices 250. Additionally or alternatively, the playback devices 250 themselves may determine their respective playback responsibilities based on the multi-channel media content received along with the input parameters 502.

[0131] In various examples, the playback responsibilities of the various playback devices 250 may be dynamically adjusted over time based on input parameters or other factors. Examples of input parameters 502 that may result in variations in playback responsibilities for one or more of the playback devices 250 include presence detection (e.g., the number of users present in a space, the distribution of people, the direction of movement, etc.), noise classification (e.g., the type and level of noise detected in the environment), time of day (e.g., circadian rhythm), or other suitable input parameters. As previously mentioned, variations in playback responsibilities can include both variations in which devices play channels and the relative levels at which particular channels are played.

[0132] FIG. 8 illustrates another example of a distributed generated media playback system 800. System 800 may be similar to system 700 described above with respect to FIG. 7, except that in system 800, local media sources 105 are coupled to group coordinator 210. This local media source 105 may be a physical line-in connection (e.g., connected to a musical instrument, microphone, record player, television, local data storage device with stored audio files, etc.) or a wireless local connection. As shown, group coordinator 210 may further include a mixer 802 configured to mix incoming media from local media sources 105 with generated media content 616 received over a network interface from a remote computing device. Performing such mixing locally may synchronize the local media content for playback with the remotely originating generated media content 616.

[0133] The mixed media content may then be transmitted from coordinator device 210 to multiple playback devices 250. As discussed above in connection with FIG. 7, the specific playback responsibilities for each of playback devices 250 may be based on (and / or may change dynamically over time based on) one or more input parameters 502, which may include sensor data, user input, metadata associated with media (whether local media or generated media content), or any other suitable input.

[0134] h. Exemplary Methods for Generating and Playing Back Generative Media 9-13 are flow diagrams of exemplary methods for playing back generated media content via multiple separate playback devices. Methods 900, 1000, 1100, 1200, and 1300 may be performed by any of the devices described herein or any other device now known or later developed.

[0135] Various examples of methods 900, 1000, 1100, 1200, and 1300 include one or more operations, functions, or actions illustrated by blocks. While the blocks are shown in sequential order, these blocks may also be performed in parallel and / or in orders different from those disclosed and described herein. Additionally, various blocks may be combined into fewer blocks, divided into additional blocks, and / or eliminated based on the desired implementation.

[0136] Furthermore, for methods 900, 1000, 1100, 1200, and 1300, as well as other processes and methods disclosed herein, flowcharts illustrate the functionality and operation of some example possible implementations. In this regard, each block may represent a module, segment, or portion of program code, including one or more instructions executable by one or more processors to implement specific logical functions or steps in the process. The program code may be stored on any type of computer-readable medium, such as a storage device including a disk or hard drive. The computer-readable medium may include non-transitory computer-readable media, such as tangible non-transitory computer-readable media for storing short-term data, such as register memory, processor cache, and random access memory (RAM). The computer-readable medium may also include non-transitory media, such as secondary or persistent long-term storage devices, such as read-only memory (ROM), optical or magnetic disks, and compact disc read-only memory (CD-ROM). The computer-readable medium may also be any other volatile or non-volatile storage system. The computer-readable medium may be considered, for example, a computer-readable storage medium or a tangible storage device. Furthermore, for the methods and other processes and methods disclosed herein, each block in Figures 9-13 may correspond to circuitry hardwired to perform specific logical functions within the process.

[0137] 9, method 900 begins at block 902 with receiving a command to play generated media content via a group of playback devices or combined zones. Such a command may be received, for example, via control device 130 or other suitable user input.

[0138] At block 904, the method 900 includes the group coordinator device providing timing information to the producing group member devices. The timing information may include contextual timing data (e.g., time data associated with sensor inputs or other user inputs), produced media playback timing data (e.g., timestamps and synchronization data to facilitate synchronized playback of the produced media), and / or media content stream timing data based on a common clock.

[0139] At block 906, the method optionally includes determining a generated media content model to be used to generate the generated media. Such a model may be implemented, for example, in the media content module 214 described above with respect to FIGS. 2-8. In some examples, each of the member devices may utilize the same or substantially the same generated media content model, while in other cases, some or all of the member devices may utilize different generated media content models. For example, a first generated media content model may generate rhythmic beats, while a second generated media content model may generate ambient nature sounds. When played simultaneously, the generated audio produced by these different generated media content models may create a pleasant listening experience for the user. In some examples, the selection of a particular generated media content model may itself be based on one or more input parameters, such as device capabilities, device location, number of users present, user sensor data, etc.

[0140] At block 908, method 900 includes the coordinator device and the member devices receiving context and / or other input data. For example, the input data may include sensor data, user input, context data, or any other relevant data that can be utilized as input for the generated media content model.

[0141] The method 900 continues at block 910 with the coordinator device and the member devices synchronizing to generate and play the generated media content.

[0142] 10 illustrates another method 1000 for playing generated audio content through multiple playback devices. Method 1000 begins with receiving one or more input parameters at a group coordinator device at block 1002. As previously mentioned, the input parameters may include sensor data, user input, contextual data, or any other input that may be used by the generated media module to generate generated audio for playback.

[0143] In block 1004, the coordinator device transmits input parameters to one or more discrete playback devices having the generated media module. For example, the coordinator device may obtain sensor data and other input parameters and transmit them to multiple discrete playback devices in the environment or to multiple discrete playback devices distributed across multiple environments. In some examples, these input parameters may include features of the generated content model itself, for example, providing instructions for updating the generated media module stored locally by one or more of the discrete playback devices.

[0144] At block 1006, the method includes transmitting timing data from the coordinator device to the playback devices. The timing data may include, for example, clock data or other synchronization signals configured to facilitate coordination of the generation of the generated media content and synchronized playback of the generated media content via the separate playback devices.

[0145] Method 1000 continues at block 1008 with simultaneously playing the generated media content via the playback devices based at least in part on the input parameters. As previously mentioned, the various playback devices may play the same generated audio, or each may play separate generated audio that, when played in sync, produces a desired psychoacoustic effect for the user present.

[0146] In the example of Figure 10, the generated media content may be generated locally by discrete playback devices, each generating and playing their own generated audio content in parallel with each other. In another method 1100 shown in Figure 11, the generated media content is generated at a coordinator device, which then sends the generated media content along with timing data to separate playback devices for synchronized playback.

[0147] At block 1102, the method 1100 includes receiving, at the group coordinator device, one or more input parameters. Examples of the input parameters are described elsewhere herein and include sensor data, user input, contextual data, or any other input that can be used by the generated media module to generate generated audio for playback.

[0148] At block 1104, the coordinator device generates first and second generated media streams based at least in part on the input parameters, and at block 906, the first and second media streams are transmitted to first and second separate playback devices, respectively. For example, the coordinator device may generate two streams forming different channels of generated audio, e.g., with a left channel played by the first playback device and a corresponding right channel played by the second playback device. Additionally or alternatively, the two streams may be separate audio tracks that can nevertheless be played synchronously, such as a rhythmic beat in one stream and ambient nature sounds in the other stream. Multiple other variations are possible. While this example describes two streams for two playback devices, in various other examples, there may be one stream or more than two streams that can be provided to any number of playback devices for synchronous playback. In at least some examples, one or more of the playback devices may be located in different environments (e.g., different homes, different cities, etc.) that are far removed from one another.

[0149] At block 1108, the first playback device plays the first originated media stream and the second playback device plays the second originated media stream. In some examples, this simultaneous playback can be facilitated by using timing data received from the coordinator device.

[0150] 12 illustrates another exemplary method 1200 for generating and playing back generated media content. As previously discussed, it may be beneficial to use one or more remote computing devices (e.g., cloud-based servers) to perform at least a portion of the processing necessary to generate the generated media content to reduce the computational demands placed on the local playback device and / or to perform operations that are not feasible using components of the local playback device. Method 1200 begins at block 1202 with receiving one or more input parameters at the playback device. As previously discussed, the input parameters may include sensor data, user input, contextual data, or any other input that may be used by the generated media module to generate generated audio for playback.

[0151] At block 1204, method 1200 includes accessing a library containing a plurality of existing media segments. For example, a plurality of individual media segments (e.g., audio tracks) may be stored on a playback device and arranged and / or mixed for playback according to a generative content model. Additionally or alternatively, the library may be stored on one or more remote computing devices, and the individual media segments may be transmitted from the remote computing devices to the playback device for playback.

[0152] The method 1000 continues at block 1206 with generating media content based at least in part on the input parameters by arranging a selection of existing media segments from a library for playback according to the generated media content model. As described elsewhere herein, the generated media content model may receive one or more input parameters as input. Based on the input, the generated media content model may be used to output particular generated media content. In examples, the generated media content may include an arrangement of existing media segments, for example, arranging them in a particular order, with or without overlap between the particular media segments, and / or with additional processing or mixing steps performed to generate the desired output.

[0153] At block 1208, the playback device plays the generated media content. In various examples, this playback can be performed simultaneously and / or synchronously with additional playback devices.

[0154] 13 illustrates an example process 1300 for playing multi-channel generated media content. As shown, method 1300 begins with receiving a stream including multiple channels of media content at a coordinator device at block 1302. For example, some or all of the channels of the multi-channel media content can be generated media.

[0155] At block 1304, method 1300 includes transmitting each of the multiple channels to multiple playback devices, including at least a first playback device and a second playback device. The coordinator device may be a playback device communicatively coupled to multiple additional playback devices in the environment. Alternatively, the coordinator device may not itself be a playback device and may send multi-channel generated media content to multiple playback devices for playback. The coordinator device may optionally divide the audio into frames for transmission to the playback devices. Additionally or alternatively, audio content may be encoded by the coordinator device for audio playback and later decoded by the playback devices.

[0156] Method 1300 continues at block 1306 with playing a first subset of channels via a first playback device according to a first playback responsibility, and at block 1308 with playing a second subset of channels via a second playback device according to a second playback responsibility. For example, multi-channel media content may include a first channel of rain sounds, a second channel of bird sounds, and a third channel of rhythmic beats. These channels may be received by each device, while only a subset of the channels is being played by any given device. For example, a first playback device may play the rain sounds and the rhythmic beats, and a second playback device may play the rain sounds and the bird sounds. Furthermore, relative levels may vary between devices. For example, a first playback device may play the rain sounds at 50% gain (i.e., gain reduced by half), and a second playback device may play the rain sounds at 100% gain (i.e., no gain reduction). By varying both the specific audio channels and the specific levels being played by the various devices, an immersive sound field can be achieved, especially in environments containing multiple playback devices working in concert.

[0157] In block 1310, the first playback responsibility and / or the second playback responsibility is dynamically changed over time. For example, playback responsibility may vary depending on which particular channels are played by different devices, the relative levels at which particular channels are played, or in other respects. In some examples, playback responsibility is changed based at least in part on one or more input parameters, such as physiological sensor data, network device sensor data, environmental data, playback device capability data, playback device state, user data, direct user input, or any other suitable input parameter. As an example, as more users enter a room, the rhythmic beat channel may be played by more playback devices than in the initial configuration.

[0158] Various examples of generated media playback are described herein. Those skilled in the art will appreciate that a wide variety of generated media modules, algorithms, inputs, sensor data, and playback device configurations are contemplated and may be used in accordance with the present technology.

[0159] IV. Conclusion The above discussion of playback devices, controller devices, playback zone configurations, and media content sources provides only a few examples of operating environments in which the features and methods described below may be implemented. Other operating environments and configurations of media playback systems, playback devices, and network devices not explicitly described herein may also be applicable and suitable for implementation of the features and methods.

[0160] The above description discloses various exemplary systems, methods, apparatus, and articles of manufacture that include, among other things, firmware and / or software executing on hardware. It is understood that such examples are merely illustrative and should not be considered limiting. For example, it is contemplated that any or all of the firmware, hardware, and / or software aspects or components could be embodied exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Thus, the examples provided are not the only ways to implement such systems, methods, apparatus, and / or articles of manufacture.

[0161] Furthermore, references herein to an "example" mean that a particular feature, structure, or characteristic described in connection with the example may be included in at least one example or embodiment of the present invention. Appearances of this phrase in various places throughout the specification do not necessarily all refer to the same example, nor are they mutually exclusive, separate, or alternative examples from other examples. Thus, it is explicitly and implicitly understood by those skilled in the art that examples described herein can be combined with other examples.

[0162] This specification is presented primarily in terms of example environments, systems, procedures, steps, logic blocks, processes, and other symbolic representations that directly or indirectly resemble the operations of network-coupled data processing devices. These process descriptions and representations are typically used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art. Numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, those skilled in the art will understand that certain examples of the present technology may be practiced without the specific specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail to avoid unnecessarily obscuring aspects of the examples. Accordingly, the scope of the present disclosure is defined by the appended claims, rather than by the foregoing description of the illustrative embodiments.

[0163] If any of the appended claims are read to cover purely software and / or firmware embodiments, at least one of the elements in at least one example is expressly defined hereby to include a tangible, non-transitory medium, such as a memory, DVD, CD, Blu-ray®, etc., that stores the software and / or firmware.

[0164] The disclosed technology is illustrated, for example, according to various examples described below. Various examples of embodiments of the disclosed technology are described as numbered examples (1, 2, 3, etc.) for convenience. These are provided as examples and are not intended to limit the disclosed technology. It should be noted that any of the subordinate examples may be combined in any combination or may be included in their own independent examples. Other examples may be presented as well.

[0165] Example 1: A method including receiving input parameters at a coordinator device; transmitting the input parameters from the coordinator device to a plurality of playback devices, each having a generated media module therein; and transmitting timing data from the coordinator device to the plurality of playback devices such that the playback devices simultaneously play the generated media content based at least in part on the input parameters.

[0166] Example 2: The method of any one of the examples herein, wherein the first and second playback devices each play different generated audio content based at least in part on the input parameters.

[0167] Example 3: The method of any one of the examples herein, wherein the input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device state (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data).

[0168] Example 4: The method of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0169] Example 5: The method of any one of the examples herein, further comprising sending a signal from the coordinator device to at least one of the plurality of playback devices that causes the playback device to modify a generated media module.

[0170] Example 6: The method of any one of the examples herein, wherein the generated media content includes at least one of generated audio content or generated visual content.

[0171] Example 7: The method of any one of the examples herein, wherein the generating media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0172] Example 8: A device comprising: a network interface; one or more processors; and a tangible, non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the device to perform operations, the operations including: receiving input parameters via the network interface; sending the input parameters via the network interface to a plurality of playback devices, each having a generated media module therein; and sending timing data to the plurality of playback devices via the network interface such that the playback devices simultaneously play the generated media content based at least in part on the input parameters.

[0173] Example 9: The device of any one of the examples herein, wherein the first and second playback devices each play different generated audio content based at least in part on the input parameters.

[0174] Example 10: The input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data) of any one of the devices in the examples herein.

[0175] Example 11: The device of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0176] Example 12: The device of any one of the examples herein, wherein the operations further include sending a signal from the coordinator device to at least one of the plurality of playback devices via a network interface, causing the generated media module of the playback device to be modified.

[0177] Example 13: The device of any one of the examples herein, wherein the generated media content includes at least one of generated audio content or generated visual content.

[0178] Example 14: The device of any one of the examples herein, wherein the generating media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0179] Example 15: A tangible, non-transitory computer-readable medium storing instructions that, when executed by one or more processors of the device, cause the device to perform operations including: receiving input parameters at a coordinator device; transmitting the input parameters from the coordinator device to a plurality of playback devices, each having a generated media module therein; and transmitting timing data from the coordinator device to the plurality of playback devices such that the playback devices simultaneously play the generated media content based at least in part on the input parameters.

[0180] Example 16: The computer-readable medium of any one of the examples herein, wherein the first and second playback devices each play different generated audio content based at least in part on the input parameters.

[0181] Example 17: The computer-readable medium of any one of the examples herein, wherein the input parameters include one or more of: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data).

[0182] Example 18: The computer-readable medium of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0183] Example 19: The computer-readable medium of any one of the examples herein, further comprising sending a signal from the coordinator device to at least one of the plurality of playback devices that causes the playback device to modify a produced media module.

[0184] Example 20: The computer-readable medium of any one of the examples herein, wherein the generated media content includes at least one of generated audio content or generated visual content.

[0185] Example 21: The computer-readable medium of any one of the examples herein, wherein the generative media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0186] Example 22: A method comprising: receiving input parameters at a coordinator device; generating first and second media content streams via a generate media module of the coordinator device; transmitting the first media content stream to a first playback device via the coordinator device; and transmitting the second media content stream to a second playback device via the coordinator device such that the first and second media content streams are played simultaneously via the first and second playback devices.

[0187] Example 23: The method of any one of the examples herein, further comprising transmitting timing data from the coordinator device to each of the first and second playback devices.

[0188] Example 24: The method of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0189] Example 25: The method of any one of the examples herein, wherein the first and second media content streams are different.

[0190] Example 26: The method of any one of the examples herein, wherein the input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data).

[0191] Example 27: The method of any one of the examples herein, further comprising modifying the originating media module of the coordinator device.

[0192] Example 28: The method of any one of the examples herein, wherein each of the first and second generated media content streams includes at least one of generated audio content or generated visual content.

[0193] Example 29: The method of any one of the examples herein, wherein the generating media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0194] Example 30: A device comprising: a network interface; a generated media module; one or more processors; and a tangible, non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the device to perform operations, the operations including receiving input parameters via the network interface; generating first and second media content streams via the generated media module; transmitting the first media content stream to a first playback device via the network interface; and transmitting the second media content stream to a second playback device via the network interface such that the first and second media content streams are played simultaneously via the first and second playback devices.

[0195] Example 31: The device of any one of the examples herein, wherein the operations further include transmitting the timing data to each of the first and second playback devices via the network interface.

[0196] Example 32: The device of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0197] Example 33: The device of any one of the examples herein, wherein the first and second media content streams are different.

[0198] Example 34: The input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data) for any one of the devices in the examples herein.

[0199] Example 35: The device of any one of the examples herein, wherein the operations further include modifying the generated media module.

[0200] Example 36: The device of any one of the examples herein, wherein each of the first and second generated media content streams includes at least one of generated audio content or generated visual content.

[0201] Example 37: The device of any one of the examples herein, wherein the generating media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0202] Example 38: A tangible, non-transitory computer-readable medium storing instructions that, when executed by one or more processors of the coordinator device, cause the coordinator device to perform operations, the operations including: receiving input parameters at the coordinator device; generating first and second media content streams via a generating media module of the coordinator device; transmitting the first media content stream to a first playback device via the coordinator device; and transmitting the second media content stream to a second playback device via the coordinator device such that the first and second media content streams are played simultaneously via the first and second playback devices.

[0203] Example 39: The computer-readable medium of any one of the examples herein, further comprising transmitting timing data from the coordinator device to each of the first and second playback devices.

[0204] Example 40: The computer-readable medium of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0205] Example 41: The computer-readable medium of any one of the examples herein, wherein the first and second media content streams are different.

[0206] Example 42: The computer-readable medium of any one of the examples herein, wherein the input parameters include one or more of: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data).

[0207] Example 43: The computer-readable medium of any one of the examples herein, wherein the operations further include modifying the generated media module of the coordinator device.

[0208] Example 44: The computer-readable medium of any one of the examples herein, wherein each of the first and second generated media content streams includes at least one of generated audio content or generated visual content.

[0209] Example 45: The computer-readable medium of any one of the examples herein, wherein the generative media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0210] Example 46: A playback device comprising: one or more amplifiers configured to drive one or more audio transducers; one or more processors; and a data storage device having instructions that, when executed by the one or more processors, cause the playback device to perform operations, the operations including: receiving one or more first input parameters at the playback device; generating first media content via the playback device based at least in part on the one or more first input parameters, the first input parameters including: accessing a library stored on the playback device that includes a plurality of pre-existing media segments, and arranging a first selection of the pre-existing media segments from the library for playback based at least in part on the one or more input parameters according to a generated media content model; and playing the first generated media content via the one or more amplifiers.

[0211] Example 47: The playback device of any one of the examples herein, wherein the operations include receiving one or more second input parameters at the playback device that differ from the first input parameters; generating second media content via the playback device based at least in part on the one or more second input parameters, where the second media content differs from the first media content, and the generating includes accessing a library and arranging a second selection of existing media segments from the library for playback based at least in part on the one or more second input parameters according to a generated media content model; and playing the second generated media content via one or more amplifiers.

[0212] Example 48: The playback device of claim 1, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments at least partially offset in time.

[0213] Example 49: The playback device of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments so that they at least partially overlap in time.

[0214] Example 50: The playback device of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes applying different equalization adjustments to different existing media segments.

[0215] Example 51: The playback device of any one of the examples herein, wherein the step of placing a first selection of existing media segments from a library or playing includes applying gain levels that vary over time to different existing media segments.

[0216] Example 52: The playback device of any one of the examples herein, wherein the step of arranging the first selection of an existing media segment from the library or playback includes randomizing a starting point for playback of the particular existing media segment.

[0217] Example 53: The playback device of any one of the examples herein, wherein the first generated media content and the second generated media content each include new media content.

[0218] Example 54: The playback device of any one of the examples herein, wherein the first generated media content includes audio content and the plurality of pre-existing media segments includes a plurality of pre-existing audio segments.

[0219] Example 55: The playback device of any one of the examples herein, wherein the first generated media content includes audio-visual content and the plurality of existing media segments includes a plurality of existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0220] Example 56: The playback device of any one of the examples herein, further including receiving additional existing media segments via the network interface and updating the library to include at least the additional existing media segments.

[0221] Example 57: The playback device of any one of the examples herein, wherein the first and second input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity, speech characteristics), user mood data).

[0222] Example 58: A method comprising: receiving one or more first input parameters at a playback device; generating first media content via the playback device based at least in part on the one or more first input parameters, the generating step comprising: accessing a library stored on the playback device that includes a plurality of pre-existing media segments, and arranging a first selection of the pre-existing media segments from the library for playback based at least in part on the one or more input parameters according to a generated media content model; and playing the first generated media content via the playback device.

[0223] Example 59: The method of any one of the examples herein, comprising: receiving one or more second input parameters at a playback device that differ from the first input parameters; generating second media content via the playback device based at least in part on the one or more second input parameters, wherein the second media content differs from the first media content, and the generating comprises accessing a library and arranging a second selection of existing media segments from the library for playback based at least in part on the one or more second input parameters according to a generated media content model; and playing the second generated media content via the playback device.

[0224] Example 60: The method of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments to be at least partially offset in time.

[0225] Example 61: The method of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments so that they are at least partially overlapping in time.

[0226] Example 62: The method of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes applying different equalization adjustments to different existing media segments.

[0227] Example 63: The method of any one of the examples herein, wherein the step of arranging a first selection of existing media segments from a library or playing includes applying gain levels that vary over time to different existing media segments.

[0228] Example 64: The method of any one of the examples herein, wherein the step of arranging the first selection of an existing media segment from the library or playing includes randomizing a starting point for playback of the particular existing media segment.

[0229] Example 65: The method of any one of the examples herein, wherein the first generated media content and the second generated media content each include new media content.

[0230] Example 66: The method of any one of the examples herein, wherein the first generated media content includes audio content and the plurality of pre-existing media segments includes a plurality of pre-existing audio segments.

[0231] Example 67: The method of any one of the examples herein, wherein the first generated media content includes audio-visual content and the plurality of existing media segments includes a plurality of existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0232] Example 68: The method of any one of the examples herein, further including receiving additional existing media segments via the network interface and updating the library to include at least the additional existing media segments.

[0233] Example 69: The method of any one of the examples herein, wherein the first and second input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity, speech characteristics), user mood data).

[0234] Example 70: A tangible, non-transitory computer-readable medium storing instructions that, when executed by one or more processors of the playback device, cause the playback device to perform operations including: receiving one or more first input parameters at the playback device; generating first media content via the playback device based at least in part on the one or more first input parameters, the generating first media content including accessing a library stored on the playback device that includes a plurality of pre-existing media segments, and arranging a first selection of the pre-existing media segments from the library for playback based at least in part on the one or more input parameters according to a generated media content model; and playing the first generated media content via the playback device.

[0235] Example 71: The computer-readable medium of any one of the examples herein, wherein the operations include receiving one or more second input parameters at a playback device that differ from the first input parameters; generating second media content via the playback device based at least in part on the one or more second input parameters, where the second media content differs from the first media content, and the generating includes accessing a library and arranging a second selection of existing media segments from the library for playback based at least in part on the one or more second input parameters according to a generated media content model; and playing the second generated media content via one or more amplifiers.

[0236] Example 72: The computer-readable medium of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments to be at least partially offset in time.

[0237] Example 73: The computer-readable medium of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments to at least partially overlap in time.

[0238] Example 74: The computer-readable medium of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes applying different equalization adjustments to different existing media segments.

[0239] Example 75: The computer-readable medium of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes applying gain levels that vary over time to different existing media segments.

[0240] Example 76: The computer-readable medium of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes randomizing a starting point for playback of the particular existing media segment.

[0241] Example 77: The computer-readable medium of any one of the examples herein, wherein the first generated media content and the second generated media content each include new media content.

[0242] Example 78: The computer-readable medium of any one of the examples herein, wherein the first generated media content includes audio content and the plurality of pre-existing media segments includes a plurality of pre-existing audio segments.

[0243] Example 79: The computer-readable medium of any one of the examples herein, wherein the first generated media content includes audio-visual content and the plurality of existing media segments includes a plurality of existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0244] Example 80: The computer-readable medium of any one of the examples herein, further including receiving additional existing media segments via the network interface and updating the library to include at least the additional existing media segments.

[0245] Example 81: The computer-readable medium of any one of the examples herein, wherein the first and second input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity, speech characteristics), user mood data).

[0246] Example 82: A method of playback comprising: a first playback device and a second playback device, the first playback device comprising: a first network interface; one or more first processors; and a data storage device having instructions having instructions that, when executed by the one or more processors, cause the first playback device to perform operations, the operations including: receiving one or more input parameters; and generating media content based at least in part on the one or more input parameters, the generated media content including a first portion and at least a second portion, the generating including: accessing a library stored on the playback device comprising a plurality of pre-existing media segments; and arranging a selection of the pre-existing media segments from the library for playback based at least in part on the one or more input parameters according to a generated media content model. , transmitting a signal including a second portion of the generated media content and corresponding timing information via a first network interface, and causing playback of the first portion of the generated media content, wherein the second playback device comprises a second network interface, one or more audio transducers, one or more second processors, and a data storage device having instructions having instructions that, when executed by the one or more second processors, cause the second playback device to perform operations, the operations including receiving the signal transmitted from the first playback device via the second network interface, and playing, via the one or more transducers, the second portion of the generated media content in accordance with the timing information substantially synchronously with the playback of the first portion of the generated media content. system.

[0247] Example 83: The system of any one of the examples herein, further comprising a network device, the network device comprising a third network interface, one or more processors, and a data storage device having instructions having instructions that, when executed by the one or more processors, cause the third playback device to perform operations, the operations including: receiving a request from the first playback device via the third network interface over a data network; and, in response to receiving the request, transmitting an updated library of existing media segments to the first playback device via the third network interface over the data network.

[0248] Example 84: The system of any one of the examples herein, wherein the network device comprises one or more of a remote server, another playback device, a mobile computing device, a laptop, or a tablet.

[0249] Example 85: A method for producing a first playback device and a second playback device communicatively coupled via a local area network, the first playback device comprising: one or more first processors; one or more first audio transducers; and a data storage device having instructions that, when executed by the one or more first processors, cause the first playback device to perform operations, the operations comprising: receiving one or more input parameters; generating first media content based at least in part on the one or more input parameters, the first playback device accessing a first library stored on the first playback device comprising a plurality of pre-existing media segments, and arranging a selection of the pre-existing media segments from the first library for playback based at least in part on the one or more input parameters according to a first produced media content model; and playing the first produced media content via the one or more first audio transducers; the second playback device comprising: a second network interface; a data storage device having one or more second audio transducers, one or more second processors, and instructions that, when executed by the one or more second processors, cause a second playback device to perform operations, the operations including: generating second media content based at least in part on one or more input parameters, the second generated media content being substantially identical to the first generated media content, the generating including accessing a second library stored on the second playback device comprising a plurality of pre-existing media segments, and arranging a selection of the pre-existing media segments from the second library for playback based at least in part on the one or more input parameters according to a second generated media content model; and playing the second generated media content via the one or more second audio transducers in synchronization with the playback of the first generated media content via the first playback device; system.

[0250] Example 86: The system of any one of the examples herein, wherein the first produced media content model and the second produced media content model are substantially identical.

[0251] Example 87: The system of any one of the examples herein, wherein the first library and the second library are substantially identical.

[0252] Example 88: A media playback system for playback of multi-channel generated media content, comprising: a first playback device comprising a first audio transducer and one or more first processors; a second playback device comprising a second audio transducer and one or more second processors; a coordinator device comprising one or more third processors; and one or more computer-readable media storing instructions that, when executed by the one or more first processors, the second processor, and / or the third processor, cause the media playback system to perform operations, wherein the operations are performed by: 1. A media playback system comprising: receiving a stream including a plurality of channels of media content, at least some of the channels including originated media content; transmitting each of the plurality of channels to a plurality of playback devices including at least a first playback device and a second playback device; playing a first subset of the channels via the first playback device according to a first playback responsibility; playing a second subset of the channels via the second playback device according to a second playback responsibility; and dynamically changing the first and / or second playback responsibilities over time.

[0253] Example 89: The system of any one of the examples herein, wherein the first playback device synchronously plays the first channel and the second channel, and the step of changing the first playback responsibility includes the step of changing the gain of playback of the first channel without changing the gain of playback of the second channel.

[0254] Example 90: The system of any one of the examples herein, wherein the dynamically changing step is based on one or more input parameters, and the input parameters include one or more of physiological sensor data, network device sensor data, environmental data, playback device capability data, playback device state, or user data.

[0255] Example 91: The system of any one of the examples herein, wherein the dynamically modifying step is responsive to user input via a control device.

[0256] Example 92: The system of any one of the examples herein, wherein the operations further include, via the coordinator device, playing back a subset of the plurality of channels according to a third playing responsibility.

[0257] Example 93: The system of any one of the examples herein, wherein the generated media content is received from one or more remote computing devices comprising generated media modules.

[0258] Example 94: The system of any one of the examples herein, wherein the operations further include receiving local media content via a physical connection at a coordinator device, mixing the local media content with a stream including multiple channels of media content via the coordinator device to generate mixed media content, and transmitting the mixed media content to multiple playback devices.

[0259] Example 95: A method for multi-channel playback of generated media content, comprising: receiving at a coordinator device a stream including multiple channels of media content, at least some of the channels including the generated media content; transmitting each of the multiple channels to a plurality of playback devices including at least a first playback device and a second playback device; playing a first subset of the channels via the first playback device according to a first playback responsibility; playing a second subset of the channels via the second playback device according to a second playback responsibility; and dynamically changing the first and / or second playback responsibilities over time.

[0260] Example 96: The method of any one of the examples herein, wherein the first playback device synchronously plays the first channel and the second channel, and the step of changing the first playback responsibility includes changing the gain of playback of the first channel without changing the gain of playback of the second channel.

[0261] Example 97: The method of any one of the examples herein, wherein the dynamically varying step is based on one or more input parameters, and the input parameters include one or more of physiological sensor data, network device sensor data, environmental data, playback device capability data, playback device state, or user data.

[0262] Example 98: The method of any one of the examples herein, wherein the dynamically modifying step is responsive to user input via a control device.

[0263] Example 99: The method of any one of the examples herein, further comprising, via the coordinator device, playing a subset of the plurality of channels according to a third playing responsibility.

[0264] Example 100: The method of any one of the examples herein, wherein the generated media content is received from one or more remote computing devices comprising generated media modules.

[0265] Example 101: The method of any one of the examples herein, further comprising receiving local media content via a physical connection at a coordinator device, mixing the local media content with a stream including multiple channels of media content via the coordinator device to generate mixed media content, and transmitting the mixed media content to multiple playback devices.

[0266] Example 102: One or more tangible, non-transitory computer-readable media storing instructions that, when executed by one or more processors of the media playback system, cause the media playback system to perform operations including: receiving, at a coordinator device, a stream including multiple channels of media content, at least some of the channels including generated media content; transmitting each of the multiple channels to multiple playback devices, including at least a first playback device and a second playback device; playing a first subset of the channels via the first playback device according to a first playback responsibility; playing a second subset of the channels via a second playback device according to a second playback responsibility; and dynamically changing the first and / or second playback responsibilities over time.

[0267] Example 103: One or more computer-readable media of any one of the examples herein, wherein a first playback device synchronously plays a first channel and a second channel, and changing the first playback responsibility includes changing a gain of playback of the first channel without changing a gain of playback of the second channel.

[0268] Example 104: One or more computer-readable media of any one of the examples herein, wherein the dynamically varying step is based on one or more input parameters, the input parameters including one or more of physiological sensor data, network device sensor data, environmental data, playback device capability data, playback device state, or user data.

[0269] Example 105: One or more computer-readable media of any one of the examples herein, wherein the dynamically modifying is responsive to user input via a control device.

[0270] Example 106: One or more computer-readable media of any one of the examples herein, wherein the operations further include, via the coordinator device, playing back a subset of the plurality of channels according to a third playing responsibility.

[0271] Example 107: One or more computer-readable media of any one of the examples herein, wherein the operations further include receiving local media content via a physical connection at a coordinator device, mixing the local media content with a stream including multiple channels of media content via the coordinator device to generate mixed media content, and transmitting the mixed media content to multiple playback devices.

Claims

1. receiving input parameters at a coordinator device; transmitting the input parameters from the coordinator device to a plurality of playback devices, each having a generated media module therein; transmitting timing data from the coordinator device to the plurality of playback devices such that the playback devices simultaneously play back generated media content based at least in part on the input parameters; A method comprising:

2. The method of claim 1 , wherein the first and second playback devices each play different generated audio content based at least in part on the input parameters.

3. The method of claim 1 or 2, wherein the timing data comprises at least one of clock data or one or more synchronization signals.

4. The method of claim 1 , further comprising the step of transmitting a signal from the coordinator device to at least one of the plurality of playback devices, the signal causing the playback device to modify the generated media module.

5. The method of claim 1 , wherein the generated media content comprises at least one of generated audio content or generated visual content.

6. receiving input parameters at a coordinator device; generating first and second media content streams via a generating media module of the coordinator device; transmitting the first media content stream to a first playback device via the coordinator device; transmitting, via the coordinator device, the second media content stream to a second playback device such that the first and second media content streams are simultaneously played via the first and second playback devices; A method comprising:

7. The method of claim 6 , further comprising transmitting timing data from the coordinator device to each of the first and second playback devices.

8. The method of claim 7 , wherein the timing data includes at least one of clock data or one or more synchronization signals.

9. The method of claim 6 , wherein the first and second media content streams are different.

10. The method of claim 6 , further comprising modifying the generated media module of the coordinator device.

11. 11. The method of claim 6, wherein each of the first and second generated media content streams comprises at least one of generated audio content or generated visual content.

12. 12. The method of any one of claims 1 to 11, wherein the generative media module includes one or more algorithms that automatically generate new media output based on inputs including at least the input parameters.

13. receiving, at a coordinator device, a stream including multiple channels of media content, at least some of the channels including originated media content; transmitting each of the plurality of channels to a plurality of playback devices, including at least a first playback device and a second playback device; playing a first subset of the channels via the first playback device according to a first playback responsibility; playing a second subset of the channels via the second playback device according to a second playback responsibility; dynamically changing the first and / or second playback responsibilities over time; 1. A method for multi-channel playback of generated media content, comprising:

14. 14. The method of claim 13, wherein the first playback device plays a first channel and a second channel synchronously, and wherein changing the first playback responsibility comprises changing a gain of playback of the first channel without changing a gain of playback of the second channel.

15. 15. The method of claim 13 or 14, wherein the dynamically varying step is in response to user input via a control device.

16. The method of claim 13 , further comprising the step of playing, via the coordinator device, a subset of the plurality of channels according to a third playing responsibility.

17. 17. The method of any one of claims 13 to 16, wherein the generated media content is received from one or more remote computing devices comprising generated media modules.

18. receiving local media content at the coordinator device via a physical connection; mixing, via the coordinator device, the local media content with the stream comprising the multiple channels of media content to generate mixed media content; transmitting the mixed media content to the plurality of playback devices; 18. The method of any one of claims 13 to 17, further comprising:

19. a coordinator device, A network interface; one or more processors; a tangible, non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the device to perform the method of any one of claims 1 to 18; a coordinator device comprising:

20. receiving one or more first input parameters at a playback device; generating first media content based at least in part on one or more of the first input parameters via the playback device; wherein the generating step includes: accessing a library stored on the playback device that includes a plurality of existing media segments; arranging a first selection of existing media segments from the library for playback based at least in part on one or more of the input parameters according to a generated media content model; playing the generated first media content via the playback device; A method comprising:

21. receiving, at the playback device, one or more second input parameters different from the first input parameters; generating, via the playback device, second media content based at least in part on one or more of the second input parameters, the second media content being different from the first media content; and the generating step further comprises: accessing the library; arranging a second selection of existing media segments from the library for playback based at least in part on one or more of the second input parameters in accordance with the generated media content model; playing the generated second media content via the playback device; 21. The method of claim 20, comprising:

22. 22. The method of claim 20 or 21, wherein arranging the first selection of existing media segments from the library for playback comprises arranging two or more of the existing media segments to be at least partially offset in time or to be at least partially overlapping in time.

23. 23. The method of any one of claims 20 to 22, wherein the generated first media content and the generated second media content each comprise new media content.

24. receiving a further existing media segment via the network interface; updating the library to include at least the additional existing media segment; 24. The method of any one of claims 20 to 23, further comprising:

25. The input parameters are: physiological sensor data, network device sensor data, Environmental data, playback device capability data, Playback device status, or User data, 25. The method of any one of claims 1 to 24, comprising one or more of:

26. 26. A tangible, non-transitory computer readable medium storing instructions that, when executed by one or more processors of a device, cause the device to perform the method of any one of claims 1 to 25.

27. A playback device, comprising: one or more amplifiers configured to drive one or more audio transducers; one or more processors; a data storage device having instructions that, when executed by the one or more processors, cause the playback device to perform the method of any one of claims 20 to 25; A playback device comprising:

28. 1. A media playback system for playing multi-channel generated media content, comprising: a first playback device comprising a first audio transducer and one or more first processors; a second playback device comprising a second audio transducer and one or more second processors; a coordinator device comprising one or more third processors; one or more computer-readable media storing instructions that, when executed by the one or more first, second, and / or third processors, cause the media playback system to perform the method of any one of claims 1 to 18 and 20 to 25; and A media playback system comprising: