Generation of Digital Media Based on Blockchain Data

The media playback system addresses the limitations of previous technologies by enabling synchronized audio playback across multiple networked devices, offering a flexible and user-friendly experience for listening to media across multiple rooms.

JP2025518556AActive Publication Date: 2025-06-17SONOS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024568610
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-18
Filing Date
2023-05-09
Publication Date
2025-06-17
Estimated Expiration
2043-05-09

AI Technical Summary

Technical Problem

Until 2002, options for accessing and listening to digital audio at high volume settings were limited, and existing technologies did not provide a seamless and synchronized media playback experience across multiple networked devices.

Method used

The development of a media playback system that enables synchronized audio playback across multiple networked devices, allowing users to play different media content in each room and group rooms for synchronous playback, using a software-controlled application on a controller device.

Benefits of technology

This solution provides a flexible and user-friendly way to experience music and other media across multiple rooms, enhancing the listening experience by allowing synchronized playback and personalized audio content based on user inputs and environmental data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025518556000001_ABST
    Figure 2025518556000001_ABST
Patent Text Reader

Abstract

Generated media content (e.g., generated audio) can be dynamically generated based on various inputs that can include blockchain data. A playback device accesses blockchain data stored via a distributed ledger and generates media content based on at least a portion of the blockchain data. The playback device can access a library of existing media segments and, according to a generated media content model, place a selection of existing media segments from the library for playback based at least in part on the blockchain data. The generated media content can be played back via the playback device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Application No. 63 / 364,931, filed May 18, 2022, the entire disclosure of which is incorporated herein by reference.

[0002] This disclosure relates to consumer goods, and more particularly, to methods, systems, products, features, services, and other elements directed to media playback or some aspects thereof.

Background Art

[0003] Until 2002 when SONOS, Inc. began developing a new type of playback system, the option to access and listen to digital audio at high volume settings was limited. Thereafter, Sonos filed one of its first patent applications, titled "Method for Synchronizing Audio Playback between Multiple Networked Devices," in 2003 and began offering its first media playback system for sale in 2005. The Sonos wireless home sound system enables people to experience music from multiple sources via one or more networked playback devices. Through a software - controlled application installed on a controller (e.g., smartphone, tablet, computer, voice - input device), one can play what one desires in any room with a networked playback device. Media content (e.g., songs, podcasts, video sound) can be streamed to the playback devices so that each room with a playback device can play a corresponding different media content. Further, rooms can be grouped together for synchronous playback of the same media content and / or the same media content can be listened to synchronously in all rooms.

Brief Description of the Drawings

[0004] The features, aspects, and advantages of the technology disclosed herein can be better understood with respect to the following description, the appended claims, and the accompanying drawings listed below. As will be appreciated by those skilled in the art, the features shown in the drawings are for illustrative purposes and variations including different and / or additional features and arrangements are possible.

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 1E

Figure 1F

Figure 1G

Figure 1H

Figure 1I

Figure 1J

Figure 1K

Figure 1L

Figure 1M

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

[0005] The drawings are for the purpose of illustrating examples of the present technology, but as will be understood by those skilled in the art, the technology disclosed herein is not limited to the arrangements and / or means shown in the drawings.

Mode for Carrying Out the Invention

[0006] I. Overview Generated media content is content that is dynamically synthesized, created, and / or modified based on an algorithm, whether implemented in software or in a physical model. Generated media content can change over time based on an algorithm alone or in relation to context data (e.g., user sensor data, environmental sensor data, event data). In various examples, such generated media content can include generated audio (e.g., music, ambient sound, etc.), generated visual images (e.g., abstract visual designs that dynamically change lighting, shape, color, etc.), generated scents, generated tactile outputs (vibrations, tactile outputs, etc.), or any other suitable media content or combinations thereof. As described elsewhere herein, generated media can be generated, at least in part, via algorithms and / or non-human systems that utilize rule-based calculations to generate new media content.

[0007] Because generated media content can be dynamically changed in real time, it enables unique user experiences that are not available using conventional media playback of pre-recorded content. For example, generated audio can be endless and / or dynamic audio that changes as input to the algorithm (e.g., input parameters related to user input, sensor data, media source data, or any other suitable input data) changes. In some examples, generated audio can be used to direct a user's mood towards a desired emotional state using one or more characteristics of the generated audio that change in response to real-time measurements reflecting the user's emotional state. As used in examples of this technology, a system can provide generated audio based on a user's current and / or desired emotional state, based on a user's activity level, based on the number of users present in the environment, or based on any other suitable input parameter.

[0008] As another example, the generated audio can be created and / or modified based on one or more inputs such as the user's location, or activity, the number of users present in the room, the time, or any other input (e.g., as determined by one or more sensors or user inputs). For example, when a single user is sitting at their desk in a calm state, the media playback system can automatically generate generated audio content suitable for focused study or work, while when multiple users are present in the room in an excited state with a lot of movement, the same media playback system can automatically generate generated audio suitable for a social gathering or dance party. In various examples, the audio characteristics that can be dynamically changed to generate the generated audio can include the selection of audio samples or clips, tempo, bass / treble / midrange volume, spatial filtering of the audio output, or any other suitable audio characteristics. The audio characteristics can be changed by using different tones or sounds, the timing of the tones or sounds, and / or audio samples that may have a desired quality. In some cases, the characteristics can also be changed by filtering or modulating the playback of the content, such as equalization, phase, or reverb / delay. During the viewing experience, the audio characteristics of the generated music can be changed based on several inputs such as various user inputs such as time, geographical location, weather, or physiological inputs such as the presumed mood, collective level of activity, or heart rate.

[0009] In some cases, generated media content such as a soundscape can be associated with data stored in one or more blockchain databases and / or layers. For example, a non-fungible token (NFT) generally consists of a data file stored on a blockchain related to a digital asset or a physical asset (e.g., a piece of music or an album, a visual artwork, a literary work). For example, even though an NFT may reference a digital art piece, it is not common for the NFT itself to contain the digital art piece itself. This is because the required file size is difficult to handle (or too costly) to store in the blockchain layer. Instead, an NFT may consist of metadata (e.g., a URL or other locator) that indicates where the digital art piece is located and / or how to access the artwork. In various examples, NFTs and other data stored via a blockchain or other distributed ledger technology can be used in media content creation. For example, in some examples, an NFT functions as an input to a generative media engine, and as a result, the generated media content has characteristics that are at least partially dependent on the data of a particular NFT or other blockchain.

[0010] Some of the examples described herein can refer to functions performed by a given party such as a "user", "listener", and / or other entity, but it should be understood that this is for illustrative purposes only. The claims should not be construed as requiring an act by the actor of any such example unless explicitly required by the language of the claims themselves.

[0011] In the figures, the same reference numbers generally indicate similar and / or identical elements. To facilitate the explanation of any particular element, the most significant digit of the reference number refers to the figure in which the element is first introduced. For example, element 110a is first introduced and described with reference to FIG. 1A. Many of the details, dimensions, angles, and other features shown in the figures are merely exemplary of specific examples of the disclosed technology. Accordingly, other examples can have other details, dimensions, angles, and features without departing from the spirit or scope of the present disclosure. Further, as will be appreciated by those skilled in the art, additional examples of the various disclosed technologies can be implemented without some of the details described below.

[0012] II. Suitable Operating Environment FIG. 1A is a partial breakaway view of a media playback system 100 distributed within an environment 101 (e.g., a home). The media playback system 100 includes one or more playback devices 110 (individually identified as playback devices 110a - 110n), one or more network microphone devices ("NMDs") 120 (individually identified as NMDs 120a - 120c), and one or more control devices 130 (individually identified as control devices 130a and 130b).

[0013] As used herein, the term "playback device" can generally refer to a network device configured to receive, process, and / or output data of a media playback system. For example, a playback device can be a network device that receives and processes audio content. In some examples, a playback device includes one or more transducers or speakers powered by one or more amplifiers. However, in other examples, a playback device includes either (or neither) of a speaker and an amplifier. For example, a playback device can include one or more amplifiers configured to drive one or more speakers external to the playback device via corresponding wires or cables.

[0014] Furthermore, as used herein, the term NMD (i.e., "Network Microphone Device") can generally refer to a network device configured for audio detection. In some examples, the NMD is a stand-alone device primarily configured for audio detection. In other examples, the NMD is incorporated into (or vice versa) a playback device.

[0015] The term "control device" can generally refer to a network device configured to perform functions related to facilitating user access, control, and / or configuration of the media playback system 100.

[0016] Each of the playback devices 110 is configured to receive an audio signal or data from one or more media sources (e.g., one or more remote servers or one or more local devices) and play the received audio signal or data as audio. One or more NMDs 120 are configured to receive voice word commands, and one or more control devices 130 are configured to receive user input. In response to the received spoken commands and / or user input, the media playback system 100 can play audio via one or more of the playback devices 110. In a particular example, the playback device 110 is configured to start playing media content in response to a trigger. For example, one or more of the playback devices 110 can be configured to play a morning playlist upon detection of a related trigger condition (e.g., detection of the presence of a user in the kitchen, detection of the operation of a coffee machine). In some examples, for instance, the media playback system 100 is configured to play audio from a first playback device (e.g., playback device 110a) in synchronization with a second playback device (e.g., playback device 110b). The interaction between the playback devices 110, NMDs 120, and / or control devices 130 of the media playback system 100 configured according to various examples of the present disclosure will be described in more detail below in connection with FIGS. 1B - 1H.

[0017] In the illustrated example of FIG. 1A, the environment 101 includes a home having several rooms, spaces, and / or playback zones, (clockwise from the upper left) a master bathroom 101a, a master bedroom 101b, a second bedroom 101c, a family room or den 101d, an office 101e, a living room 101f, a dining room 101g, a kitchen 101h, and an outdoor patio 101i. Although specific examples and instances are described below in the context of a home environment, the techniques described herein may be implemented in other types of environments. In some examples, for instance, the media playback system 100 can be implemented in one or more commercial settings (such as restaurants, malls, airports, hotels, retail stores or other shops), one or more vehicles (such as sport utility vehicles, buses, cars, ships, boats, airplanes), multiple environments (such as a combination of a home environment and a vehicle environment), and / or other suitable environments where multi-zone audio may be desirable.

[0018] The media playback system 100 can comprise one or more playback zones, some of which may correspond to rooms within the environment 101. The media playback system 100 can be established using one or more playback zones, for example, after additional zones can be added or removed to form the configuration shown in FIG. 1A. Each zone can be named according to a different room or space, such as the office 101e, the master bathroom 101a, the master bedroom 101b, the second bedroom 101c, the kitchen 101h, the dining room 101g, the living room 101f, and / or the outdoor patio 101i. In some aspects, a single playback zone can include multiple rooms or spaces. In certain aspects, a single room or space can include multiple playback zones.

[0019] In the illustrated example of FIG. 1A, the master bathroom 101a, the second bedroom 101c, the office 101e, the living room 101f, the dining room 101g, the kitchen 101h, and the outdoor patio 101i each include one playback device 110, and the master bedroom 101b and the private room 101d include a plurality of playback devices 110. In the master bedroom 101b, the playback devices 110l and 110m may be configured to play audio content synchronously, for example, as individual ones of the playback devices 110, as a combined playback zone, as an integrated playback device, and / or as any combination thereof. Similarly, in the private room 101d, the playback devices 110h to 110j may be configured to play audio content synchronously, for example, as individual devices of the playback devices 110, as one or more combined playback devices, and / or as one or more integrated playback devices. Further details regarding combined and integrated playback devices are described below with respect to FIGS. 1B and 1E.

[0020] In some aspects, one or more of the playback zones within environment 101 may each be playing different audio content. For example, a user may be grilling in patio 101i and listening to hip-hop music being played by playback device 110c, while another user may be preparing food in kitchen 101h and listening to classical music being played by playback device 110b. In another example, the playback zones can play the same audio content in synchronization with another playback zone. For example, a user in office 101e can hear that playback device 110f is playing the same hip-hop music being played by playback device 110c on patio 101i. In some aspects, playback devices 110c and 110f synchronously play hip-hop music such that the user perceives that the audio content is being played seamlessly (or at least substantially seamlessly) as the playback devices move between different playback zones. Further details regarding audio playback synchronization between playback devices and / or zones can be found, for example, in U.S. Patent No. 8,234,395, entitled "System and Method for Synchronizing Operations Among Multiple Independently Clock-Controlled Digital Data Processing Devices," which is hereby incorporated by reference in its entirety.

[0021] a. Suitable Media Playback System FIG. 1B is a schematic diagram of media playback system 100 and cloud network 102. For ease of illustration, certain devices of media playback system 100 and cloud network 102 are omitted from FIG. 1B. One or more communication links 103 (hereinafter referred to as "link 103") communicatively couple media playback system 100 and cloud network 102.

[0022] Link 103 can include, for example, one or more wired networks, one or more wireless networks, one or more wide area networks (WANs), one or more local area networks (LANs), one or more personal area networks (PANs), one or more telecommunications networks (e.g., one or more Global System for Mobile (GSM) networks, Code Division Multiple Access (CDMA) networks, Long Term Evolution (LTE) networks, 5G communication networks, and / or other suitable data transmission protocol networks). Cloud network 102 is configured to deliver media content (e.g., audio content, video content, photos, social media content) to media playback system 100 in response to requests transmitted from media playback system 100 via link 103. In some examples, cloud network 102 is further configured to receive data (e.g., voice input data) from media playback system 100 and, in response, transmit commands and / or media content to media playback system 100.

[0023] The cloud network 102 comprises computing devices 106 (identified separately as the first computing device 106a, the second computing device 106b, and the third computing device 106c). The computing devices 106 can comprise individual computers or servers such as, for example, media streaming service servers that store audio and / or other media content, voice service servers, social media servers, media playback system control servers, etc. In some examples, one or more of the computing devices 106 comprise modules of a single computer or server. In a particular example, one or more of the computing devices 106 comprise one or more modules, computers, and / or servers. Further, although the cloud network 102 is described above in the context of a single cloud network, in some examples, the cloud network 102 comprises a plurality of cloud networks including communicatively coupled computing devices. Further, although the cloud network 102 is shown in FIG. 1B as having three of the computing devices 106, in some examples, the cloud network 102 comprises fewer (or more) than three computing devices 106.

[0024] The media playback system 100 is configured to receive media content from the network 102 via the link 103. The received media content can include, for example, a Uniform Resource Identifier (URI) and / or a Uniform Resource Locator (URL). For example, in some instances, the media playback system 100 can stream, download, or obtain data from the URI or URL corresponding to the received media content. The network 104 communicatively couples the link 103 with at least a portion of the devices of the media playback system 100 (e.g., one or more of the playback device 110, the NMD 120, and / or the control device 130). The network 104 can include, for example, a wireless network (e.g., a WiFi network, Bluetooth, a Z-Wave network, ZigBee, and / or another suitable wireless communication protocol network) and / or a wired network (e.g., Ethernet, Universal Serial Bus (USB), and / or a network including another suitable wired communication). As will be understood by those skilled in the art, as used herein, "WiFi" can refer to several different communication protocols, such as, for example, Institute of Electrical and Electronics Engineers (IEEE) 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.11ac, 802.11ad, 802.11af, 802.11ah, 802.11ai, 802.11aj, 802.11aq, 802.11ax, 802.11ay, 802.15, etc., transmitted at 2.4 gigahertz (GHz), 5 GHz, and / or another suitable frequency.

[0025] In some examples, the network 104 comprises a dedicated communication network for the media playback system 100 to send messages between individual devices and / or to send media content between the media content source (e.g., one or more of the computing devices 106). In a particular example, the network 104 is configured to be accessible only to the devices within the media playback system 100, thereby reducing interference and contention with other household devices. However, in other examples, the network 104 comprises an existing household communication network (e.g., a household WiFi network). In some examples, the link 103 and the network 104 comprise one or more of the same network. In some aspects, for example, the link 103 and the network 104 comprise a telecommunications network (e.g., an LTE network, a 5G network). Further, in some examples, the media playback system 100 is implemented without the network 104, and the devices comprising the media playback system 100 can communicate with each other via, for example, one or more direct connections, a PAN, a telecommunications network, and / or other suitable communication links.

[0026] In some examples, the audio content source may be periodically added or removed from the media playback system 100. In some examples, for example, the media playback system 100 performs indexing of media items when one or more media content sources are updated, added, and / or removed from the media playback system 100. The media playback system 100 can scan for identifiable media items within some or all of the folders and / or directories accessible to the playback device 110 and generate or update a media content database with metadata (e.g., title, artist, album, track length) and other relevant information (e.g., URI, URL) for each of the discovered identifiable media items. In some examples, for example, the media content database is stored on one or more of the playback device 110, the NMD 120, and / or the control device 130.

[0027] In the illustrated example of FIG. 1B, playback devices 110l and 110m comprise group 107a. Playback devices 110l and 110m are located in different rooms in the home and can be grouped together in group 107a temporarily or permanently based on user input received at control devices 130a and / or another control device 130 within media playback system 100. When arranged in group 107a, playback devices 110l and 110m can be configured to synchronously play the same or similar audio content from one or more audio content sources. In a particular example, for instance, group 107a includes a combined zone in which playback devices 110l and 110m each include a left audio channel and a right audio channel of multi-channel audio content, thereby generating or enhancing the stereo effect of the audio content. In some examples, group 107a includes additional playback devices 110. However, in other examples, media playback system 100 omits group 107a and / or other grouped arrangements of playback devices 110.

[0028] The media playback system 100 includes NMD120a and 120d, each of which is equipped with one or more microphones configured to receive voice utterances from a user. In the example illustrated in FIG. 1B, NMD120a is a stand-alone device, and NMD120d is incorporated into the playback device 110n. NMD120a is configured to receive, for example, voice input 121 from user 123. In some examples, NMD120a transmits data related to the received voice input 121 to a voice assistant service (VAS) configured to (i) process the received voice input data and (ii) send corresponding commands to the media playback system 100. In some aspects, for example, the computing device 106c includes one or more modules and / or servers of a VAS (e.g., SONOS®, AMAZON®, GOOGLE®, APPLE®, MICROSOFT®). The computing device 106c can receive voice input data from NMD120a via the network 104 and the link 103. In response to receiving the voice input data, the computing device 106c processes the voice input data (i.e., "Play the Beatles' Hey Jude song") and determines that the processed voice input includes a command to play a song (e.g., "Hey Jude"). Accordingly, the computing device 106c sends a command to the media playback system 100 to play "Hey Jude" by the Beatles from one or more appropriate media services of the playback devices 110 (e.g., via one or more of the computing devices 106).

[0029] b. Appropriate playback device FIG. 1C is a block diagram of a playback device 110a having an input / output 111. The input / output 111 can include an analog I / O 111a (e.g., one or more wires, cables, and / or other suitable communication links configured to carry analog signals) and / or a digital I / O 111b (e.g., one or more wires, cables, or other suitable communication links configured to carry digital signals). In some examples, the analog I / O 111a is an audio line input connection that includes, for example, an auto-detect 3.5 mm audio line input connection. In some examples, the digital I / O 111b comprises a Sony / Philips Digital Interface Format (S / PDIF) communication interface and / or cable and / or a Toshiba Link (TOSLINK) cable. In some examples, the digital I / O 111b comprises a High-Definition Multimedia Interface (HDMI (registered trademark)) interface and / or cable. In some examples, the digital I / O 111b includes one or more wireless communication links that include, for example, radio frequency (RF), infrared, WiFi, Bluetooth, or another suitable communication protocol. In a particular example, the analog I / O 111a and the digital 111b each comprise an interface (e.g., a port, plug, jack) configured to receive a connector of a cable that transmits an analog signal and a digital signal, respectively, without necessarily including a cable.

[0030] The playback device 110a can receive media content (e.g., audio content including music and / or other sounds) from the local audio source 105 via, for example, an input / output 111 (e.g., a cable, wire, PAN, Bluetooth connection, ad-hoc wired or wireless communication network, and / or another suitable communication link). The local audio source 105 can comprise, for example, a mobile device (e.g., a smartphone, tablet, laptop computer) or another suitable audio component (e.g., a television, desktop computer, amplifier, phonograph, Blu-ray player, memory storing digital media files). In some aspects, the local audio source 105 includes a local music library on a smartphone, computer, network-attached storage (NAS), and / or another suitable device configured to store media files. In certain examples, one or more of the playback device 110, NMD 120, and / or control device 130 comprises the local audio source 105. However, in other examples, the media playback system completely omits the local audio source 105. In some examples, the playback device 110a does not include the input / output 111 and receives all audio content via the network 104.

[0031] The playback device 110a further includes an electronic device 112, a user interface 113 (e.g., one or more buttons, knobs, dials, touch-sensitive surfaces, displays, touchscreens), and one or more transducers 114 (hereinafter referred to as "transducer 114"). The electronic device 112 receives audio from an audio source (e.g., local audio source 105) via the input / output 111, or from one or more of the computing devices 106a - 106c via the network 104 (FIG. 1B), amplifies the received audio, and outputs the amplified audio for playback via one or more of the transducers 114. In some examples, the playback device 110a optionally includes one or more microphones 115 (e.g., a single microphone, multiple microphones, microphone array) (hereinafter referred to as "microphone 115"). In a particular example, for instance, the playback device 110a having one or more of the optional microphones 115 can operate as an NMD configured to receive voice input from a user and perform corresponding one or more operations based on the received voice input.

[0032] In the illustrated example of FIG. 1C, the electronic device 112 includes one or more processors 112a (hereinafter referred to as "processor 112a"), a memory 112b, software components 112c, a network interface 112d, one or more audio processing components 112g (hereinafter referred to as "audio component 112g"), one or more audio amplifiers 112h (hereinafter referred to as "amplifier 112h"), and a power source 112i (e.g., one or more power supplies, power cables, power receptacles, batteries, induction coils, Power over Ethernet (POE) interfaces, and / or other suitable power sources). In some embodiments, the electronic device 112 optionally includes one or more other components 112j (e.g., one or more sensors, video displays, touchscreens, battery charging bases).

[0033] Processor 112a can include a clock-driven computing component configured to process data, and memory 112b can include a computer-readable medium (e.g., a tangible non-transitory computer-readable medium loaded with one or more of software components 112c, a data storage device) configured to store instructions for performing various operations and / or functions. Processor 112a is configured to execute instructions stored in memory 112b to perform one or more of the operations. The operations can include, for example, causing playback device 110a to retrieve audio data from an audio source (e.g., one or more of computing devices 106a-106c (FIG. 1B)) and / or another one of playback devices 110. In some examples, the operations can further include causing playback device 110a to transmit audio data to another device of playback device 110a and / or another device (e.g., one of NMDs 120). A particular example includes an operation of causing playback device 110a to pair with another one of one or more playback devices 110 to enable a multi-channel audio environment (e.g., a stereo pair, a combined zone).

[0034] Processor 112a can be further configured to execute an operation of synchronizing playback of audio content on playback device 110a with another one of one or more playback devices 110. As would be understood by one of ordinary skill in the art, during synchronous playback of audio content on multiple playback devices, a listener preferably cannot perceive a time delay difference between playback of the audio content by playback device 110a and playback of the audio content by one or more other playback devices 110. Further details regarding audio playback synchronization between playback devices can be found, for example, in U.S. Patent No. 8,234,395, which is incorporated by reference above.

[0035] In some examples, memory 112b is further configured to store data associated with playback device 110a, such as one or more zones and / or zone groups of which playback device 110a is a member, audio sources accessible to playback device 110a, and / or a playback queue that playback device 110a (and / or other playback devices of the one or more playback devices) may be associated with. The stored data can be updated periodically and can include one or more state variables used to describe the state of playback device 110a. Memory 112b can also include data associated with the state of one or more of the other devices of media playback system 100 (e.g., playback device 110, NMD 120, control device 130). In some aspects, for example, the state data is shared among at least some of the devices of media playback system 100 at a predetermined time interval (e.g., every 5 seconds, every 10 seconds, every 60 seconds), such that one or more of the devices have the latest data associated with media playback system 100.

[0036] Network interface 112d is configured to facilitate the transmission of data between playback device 110a and one or more other devices on a data network such as, for example, link 103 and / or network 104 (FIG. 1B). Network interface 112d is configured to transmit and receive data corresponding to other signals (e.g., non-transitory signals) including media content (e.g., audio content, video content, text, photos), and digital packet data including an Internet Protocol (IP)-based source address and / or an IP-based destination address. Network interface 112d can analyze the digital packet data such that electronic device 112 can properly receive and process data destined for playback device 110a.

[0037] In the illustrated example of FIG. 1C, network interface 112d includes one or more wireless interfaces 112e (hereinafter referred to as "wireless interface 112e"). The wireless interface 112e (e.g., a suitable interface including one or more antennas) can be configured to wirelessly communicate with one or more other devices (e.g., one or more of other playback devices 110, NMD 120, and / or control device 130) communicatively coupled to network 104 (FIG. 1B) according to a suitable wireless communication protocol (e.g., WiFi, Bluetooth, LTE). In some examples, network interface 112d optionally includes a wired interface 112f (e.g., an interface or receptacle configured to receive a network cable such as Ethernet, USB-A, USB-C, and / or Thunderbolt cable) configured to communicate with other devices via a wired connection according to a suitable wired communication protocol. In a particular example, network interface 112d includes wired interface 112f and excludes wireless interface 112e. In some examples, electronic device 112 completely excludes network interface 112d and transmits and receives media content and / or other data via another communication path (e.g., input / output 111).

[0038] The audio component 112g is configured to process and / or filter data including media content received by the electronic device 112 (e.g., via the input / output 111 and / or the network interface 112d) to generate an output audio signal. In some examples, the audio processing component 112g includes, for example, one or more digital-to-analog converters (DACs), audio preprocessing components, audio enhancement components, digital signal processors (DSPs), and / or other suitable audio processing components, modules, circuits, etc. In a particular example, one or more of the audio processing components 112g can include one or more sub-components of the processor 112a. In some examples, the electronic device 112 omits the audio processing component 112g. In some aspects, for example, the processor 112a executes instructions stored in the memory 112b to perform audio processing operations to generate an output audio signal.

[0039] Amplifier 112h is configured to receive and amplify an audio output signal generated by audio processing component 112g and / or processor 112a. Amplifier 112h can comprise an electronic device and / or component configured to amplify the audio signal to a level sufficient to drive one or more of transducers 114. In some examples, for instance, amplifier 112h includes one or more switching or class-D power amplifiers. However, in other examples, the amplifier includes one or more other types of power amplifiers (e.g., linear gain power amplifier, class-A amplifier, class-B amplifier, class-AB amplifier, class-C amplifier, class-D amplifier, class-E amplifier, class-F amplifier, class-G amplifier, and / or class-H amplifier, and / or another suitable type of power amplifier). In a particular example, amplifier 112h comprises a suitable combination of two or more of the aforementioned types of power amplifiers. Further, in some examples, individual ones of amplifier 112h correspond to individual ones of transducers 114. However, in other examples, electronic device 112 includes a single amplifier of amplifier 112h configured to output the amplified audio signal to a plurality of transducers 114. In some other examples, electronic device 112 omits amplifier 112h.

[0040] The transducer 114 (e.g., one or more speakers and / or speaker drivers) receives the amplified audio signal from the amplifier 112h and renders or outputs the amplified audio signal as sound (e.g., audible sound waves having frequencies from about 20 Hertz (Hz) to 20 kilohertz (kHz)). In some examples, the transducer 114 can comprise a single transducer. However, in other examples, the transducer 114 comprises a plurality of audio transducers. In some examples, the transducer 114 comprises a plurality of types of transducers. For example, the transducer 114 can include one or more low-frequency transducers (e.g., subwoofers, woofers), mid-frequency transducers (e.g., midrange transducers, midwoofers), and one or more high-frequency transducers (e.g., one or more tweeters). As used herein, "low frequency" can generally refer to audible frequencies below about 500 Hz, "mid-frequency" can generally refer to audible frequencies from about 500 Hz to about 2 kHz, and "high frequency" can generally refer to audible frequencies above 2 kHz. However, in certain examples, one or more of the transducers 114 comprises a transducer that does not adhere to the aforementioned frequency ranges. For example, one of the transducers 114 may comprise an intermediate woofer transducer configured to output sound at frequencies from about 200 Hz to about 5 kHz.

[0041] By way of example, SONOS, Inc. currently offers (or offers) for sale certain playback devices including, for example, "SONOS ONE", "PLAY:1", "PLAY:3", "PLAY:5", "PLAYBAR", "PLAYBASE", "CONNECT:AMP", "CONNECT", and "SUB". In addition to or in place of this, other suitable playback devices may be used to implement the playback devices of the examples disclosed herein. Further, as will be understood by those skilled in the art, the playback devices are not limited to the examples described herein or the SONOS product offerings. In some examples, for instance, one or more playback devices 110 include wired or wireless headphones (e.g., over-ear headphones, on-ear headphones, in-ear earphones). In other examples, one or more of the playback devices 110 include a docking station and / or interface configured to interact with a docking station for a personal mobile media playback device. In certain examples, the playback device may be integrated with another device or component such as a television, lighting fixture, or some other device for use indoors or outdoors. In some examples, the playback device omits the user interface and / or one or more transducers. For example, FIG. 1D is a block diagram of a playback device 110p that includes an input / output 111 and an electronic device 112 that do not have a user interface 113 or a transducer 114.

[0042] FIG. 1E is a block diagram of a combined playback device 110q comprising a playback device 110a (FIG. 1C) ultrasonically bonded to a playback device 110i (e.g., a subwoofer) (FIG. 1A). In the illustrated example, playback devices 110a and 110i are separate ones of the playback devices 110 housed in separate enclosures. However, in some examples, the combined playback device 110q comprises a single enclosure that houses both playback devices 110a and 110i. The combined playback device 110q can be configured to process and play sound in a different manner than an uncombined playback device (e.g., playback device 110a of FIG. 1C) and / or a paired or combined playback device (e.g., playback devices 110l and 110m of FIG. 1B). In some examples, for instance, playback device 110a is a full-range playback device configured to render low-frequency, mid-frequency, and high-frequency audio content, and playback device 110i is a subwoofer configured to render low-frequency audio content. In some aspects, playback device 110a is configured to render only the mid-frequency and high-frequency components of a particular audio content when combined with a first playback device, while playback device 110i renders the low-frequency component of the particular audio content. In some examples, the combined playback device 110q includes additional playback devices and / or another combined playback device.

[0043] c. Appropriate Network Microphone Device (NMD) Figure 1F is a block diagram of NMD120a (FIGS. 1A and 1B). NMD120a includes one or more audio processing components 124 (hereinafter, “audio component 124”) and some of the components described with respect to playback device 110a (FIG. 1C) including processor 112a, memory 112b, and microphone 115. NMD120a optionally includes other components also included in playback device 110a (FIG. 1C) such as user interface 113 and / or transducer 114. In some examples, NMD120a is configured as a media playback device (e.g., one or more of playback devices 110) and further includes, for example, one or more of audio component 112g (FIG. 1C), amplifier 114, and / or other playback device components. In a particular example, NMD120a comprises an Internet of Things (IoT) device such as, for example, a thermostat, an alarm panel, a fire and / or smoke detector. In some examples, NMD120a includes microphone 115, audio processing 124, and only some of the components of electronic device 112 described above with respect to FIG. 1B. In some aspects, for example, NMD120a includes processor 112a and memory 112b (FIG. 1B) while omitting one or more other components of electronic device 112. In some examples, NMD120a includes additional components (e.g., one or more sensors, cameras, thermometers, barometers, hygrometers).

[0044] In some examples, the NMD can be incorporated into a playback device. FIG. 1G is a block diagram of a playback device 110r that includes an NMD 120d. The playback device 110r can include many or all of the components of the playback device 110a and can further include a microphone 115 and audio processing 124 (FIG. 1F). The playback device 110r may have an integrated control device 130c. The control device 130c can include, for example, a user interface (e.g., the user interface 113 of FIG. 1B) configured to receive user input (e.g., touch input, voice input) without the accompaniment of a separate control device. However, in other examples, the playback device 110r receives commands from another control device (e.g., the control device 130a of FIG. 1B).

[0045] Referring again to FIG. 1F, the microphone 115 is configured to acquire, capture, and / or receive sound from the environment in which the NMD 120a is located (e.g., the environment 101 of FIG. 1A) and / or the room. The received audio can include, for example, voice utterances, audio reproduced by the NMD 120a and / or another playback device, background audio, ambient sound, and the like. The microphone 115 converts the received sound into an electrical signal to generate microphone data. The audio processing 124 receives and analyzes the microphone data to determine whether there is a voice input in the microphone data. The voice input can include, for example, an activation word followed by an utterance that includes a user request. As will be understood by those skilled in the art, the activation word is a word or other audio cue that means a user voice input. For example, when querying AMAZON (registered trademark) VAS, a user may utter the activation word "Alexa". Other examples include "OK, Google" for invoking Google (registered trademark) VAS and "Hey, Siri" for invoking Apple (registered trademark) VAS.

[0046] After detecting the wake word, the voice processing 124 monitors the microphone data for user requests associated with the voice input. The user requests can include, for example, commands for controlling third-party devices such as a thermostat (e.g., NEST (registered trademark) thermostat), a lighting device (e.g., PHILIPS HUE (registered trademark) lighting device), or a media playback device (e.g., Sonos (registered trademark) playback device). For example, the user can set the temperature in the home by speaking the utterance "Set the thermostat to 68 degrees" following the wake word "Alexa" (e.g., in the environment 101 of FIG. 1A). The user can turn on the lighting devices in the living area of the house by speaking the utterance "Turn on the living room" after uttering the same wake word. The user can similarly speak a wake word followed by a request to play a specific song, album, or music playlist on a playback device in the home.

[0047] d. Appropriate control device FIG. 1H is a partial schematic view of a control device 130a (FIGS. 1A and 1B). As used herein, the term "control device" can be used interchangeably with "controller" or "control system". Among other features, the control device 130a is configured to receive user input related to the media playback system 100 and, in response, cause one or more devices within the media playback system 100 to perform an operation or operations corresponding to the user input. In the illustrated example, the control device 130a comprises a smartphone (e.g., iPhone (registered trademark), Android phone) on which media playback system controller application software is installed. In some examples, the control device 130a comprises, for example, a tablet (e.g., iPad (registered trademark)), a computer (e.g., a laptop computer, a desktop computer), and / or another suitable device (e.g., a television, an automotive audio head unit, an IoT device). In a particular example, the control device 130a comprises a dedicated controller for the media playback system 100. In other examples, as previously described with respect to FIG. 1G, the control device 130a is incorporated into another device within the media playback system 100 (e.g., a playback device 110, an NMD 120, and / or one or more of the other suitable devices configured to communicate via a network).

[0048] The control device 130a includes an electronic device 132, a user interface 133, one or more speakers 134, and one or more microphones 135. The electronic device 132 includes one or more processors 132a (hereinafter referred to as "processor 132a"), a memory 132b, software components 132c, and a network interface 132d. The processor 132a can be configured to execute functions related to facilitating user access, control, and configuration of the media playback system 100. The memory 132b can include a data storage device capable of loading one or more of the software components executable by the processor 112a to execute these functions. The software components 132c can include applications and / or other executable software configured to facilitate control of the media playback system 100. The memory 112b can be configured to store, for example, the software components 132c, media playback system controller application software, and / or other data related to the media playback system 100 and the user.

[0049] Network interface 132d is configured to facilitate network communication between control device 130a and one or more other devices within media playback system 100 and / or one or more remote devices. In some examples, network interface 132d is configured to operate in accordance with one or more suitable communication industry standards (e.g., infrared, wireless, wired standards including IEEE802.3, wireless standards including IEEE802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G, LTE). Network interface 132d can be configured to transmit data to and / or receive data from, for example, one of playback device 110, NMD 120, other control device 130, one of computing devices 106 of FIG. 1B, a device comprising one or more other media playback systems, etc. The data transmitted and / or received can include, for example, playback device control commands, state variables, playback zones, and / or zone group configurations. For example, based on user input received at user interface 133, network interface 132d can transmit playback device control commands (e.g., volume control, audio playback control, audio content selection) from control device 130 to one or more of playback devices 110. Network interface 132d can also transmit and / or receive configuration changes such as, for example, addition / removal of one or more playback devices 110 to / from a zone, addition / removal of one or more zones to / from a zone group, formation of combined or integrated players, separation of one or more playback devices from a combined or integrated player. An explanation of zone and group addition can be found below with respect to FIGS. 1I - 1M.

[0050] The user interface 133 is configured to receive user input and can facilitate the control of the media playback system 100. The user interface 133 includes media content technology 133a (e.g., album art, lyrics, video), a playback status indicator 133b (e.g., elapsed time and / or remaining time indicator), a media content information area 133c, a playback control area 133d, and a zone indicator 133e. The media content information area 133c can include the display of the currently playing media content and / or related information (e.g., title, artist, album, genre, release year) regarding the media content in the queue or playlist. The playback control area 133d includes selectable (e.g., via touch input and / or via a cursor or another suitable selector) icons for causing one or more playback devices in the selected playback zone or zone group to perform playback actions such as play or pause, fast forward, rewind, skip to next, skip to previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit crossfade mode, etc. The playback control area 133d can also include selectable icons for changing equalization settings, playback volume, and / or other suitable playback operations. In the illustrated example, the user interface 133 comprises a display presented on the touch screen interface of a smartphone (e.g., iPhone (registered trademark), Android phone). However, in some examples, one or more network devices can alternatively implement user interfaces of various formats, styles, and interactive sequences to provide equivalent control access to the media playback system.

[0051] One or more speakers 134 (e.g., one or more transducers) may be configured to output sound to a user of the control device 130a. In some examples, the one or more speakers comprise individual transducers configured to output corresponding low frequencies, midrange frequencies, and / or high frequencies. In some aspects, for example, the control device 130a is configured as a playback device (e.g., one of the playback devices 110). Similarly, in some examples, the control device 130a is configured as an NMD (e.g., one of the NMDs 120) and receives voice commands and other sounds via one or more microphones 135.

[0052] One or more microphones 135 can comprise, for example, one or more condenser microphones, electret condenser microphones, dynamic microphones, and / or other suitable types of microphones or transducers. In some examples, two or more of the microphones 135 are arranged to capture position information of an audio source (e.g., voice, audible sound) and / or are configured to facilitate filtering of background noise. Further, in certain examples, the control device 130a is configured to operate as a playback device and an NMD. However, in other examples, the control device 130a omits one or more speakers 134 and / or one or more microphones 135. For example, the control device 130a can comprise a device (e.g., a thermostat, an IoT device, a network device) that is part of the electronic device 132 and includes a user interface 133 (e.g., a touch screen) without a speaker or microphone.

[0053] Appropriate playback device configuration Figures 1I to 1M illustrate exemplary configurations of playback devices in zones and zone groups. First, referring to Figure 1M, in one example, a single playback device can belong to a zone. For example, the playback device 110g in the second bedroom 101c (Figure 1A) may belong to Zone C. In some implementations described below, a plurality of playback devices can be "combined" to form a "combination pair", which together form a single zone. For example, the playback device 110l (e.g., the left playback device) can be combined with the playback device 110l (e.g., the left playback device) to form Zone A. The combined playback devices may have different playback responsibilities (e.g., channel responsibilities). In another embodiment described below, a plurality of playback devices can be merged to form a single zone. For example, the playback device 110h (e.g., the front playback device) may be merged with the playback device 110i (e.g., the subwoofer) and the playback devices 110j and 110k (e.g., the left and right surround speakers respectively) to form a single Zone D. In another example, the playback devices 110g and 110h can be merged to form a merged group or zone group 108b. The merged playback devices 110g and 110h may not be specifically assigned different playback responsibilities. That is, the merged playback devices 110h and 110i can play audio content in the same way as if they were not merged, apart from playing audio content synchronously.

[0054] Each zone within the media playback system 100 may be provided for control as a single user interface (UI) entity. For example, Zone A may be provided as a single entity called the master bathroom. Zone B may be provided as a single entity called the master bedroom. Zone C may be provided as a single entity called the second bedroom.

[0055] The combined playback devices can have different playback responsibilities, such as responsibility for a specific audio channel. For example, as shown in FIG. 1-I, the playback devices 110l and 110m may be combined to generate or enhance the stereo effect of the audio content. In this example, the playback device 110l may be configured to play the audio components of the left channel, while the playback device 110k may be configured to play the audio components of the right channel. In some implementations, such a stereo combination may be referred to as "pairing".

[0056] Furthermore, the combined playback devices may have additional and / or different respective speaker drivers. As shown in FIG. 1J, the playback device 110h named Front may be combined with the playback device 110i named SUB. The front device 110h can be configured to render the mid-frequency to high-frequency range, and the SUB device 110i can be configured to render the low frequencies. However, when not combined, the front device 110h can be configured to render the entire frequency range. As another example, FIG. 1K shows the front device 110h and the SUB device 110i further combined with the left playback device 110j and the right playback device 110k respectively. In some implementations, the right device 110j and the left device 102k can be configured to form the surround or "satellite" channels of a home theater system. The combined playback devices 110h, 110i, 110j, and 110k can form a single zone D (FIG. 1M).

[0057] The merged playback device may not be assigned a playback responsibility, and each playback device can render the entire range of audio content that each is capable of. Nevertheless, the merged device may be represented as a single UI entity (i.e., a zone as described above). For example, the playback devices 110a and 110n in the master bathroom have a single UI entity for Zone A. In one example, the playback devices 110a and 110n can each output the entire range of audio content that the respective playback devices 110a and 110n can synchronize to.

[0058] In some examples, the NMD is combined or merged with another device to form a zone. For example, the NMD 120b may be combined with the playback device 110e that together form Zone F, which is called the living room. In other examples, a stand-alone network microphone device may itself be within a zone. However, in other examples, a stand-alone network microphone device may not be associated with a zone. Further details regarding associating a network microphone device and a playback device as a designated device or a default device can be found, for example, in U.S. Patent Application No. 15 / 438,749, referenced above.

[0059] Zones of individual devices, combined devices, and / or merged devices can be grouped to form zone groups. For example, referring to FIG. 1M, zone A can be grouped with zone B to form zone group 108a that includes two zones. Similarly, zone G can be grouped with zone H to form zone group 108b. As another example, zone A may be grouped with one or more other zones C-I. Zones A-I can be grouped and ungrouped in a number of ways. For example, three, four, five, or more (e.g., all) of zones A-I may be grouped. When grouped, zones of individual and / or combined playback devices can play audio in synchronization with each other, as described in the previously referenced U.S. Patent No. 8,234,395. Playback devices may be dynamically grouped and ungrouped to form new or different groups that play audio content in synchronization.

[0060] In various embodiments, the zones within the environment may be a combination of the default names of the zones within the group or the names of the zones within the zone group. For example, zone group 108b can be assigned a name such as "Dining + Kitchen" as shown in FIG. 1M. In some examples, the zone group may be given a unique name selected by the user.

[0061] Certain data may be stored in the memory of a playback device (e.g., memory 112b of FIG. 1C) as one or more state variables that are periodically updated and used to describe the state of the playback zone, the playback device, and / or the zone group associated with the playback zone. The memory may also be associated with the states of other devices in the media system and can include data that is sometimes shared between devices so that one or more of the devices have the most recent data associated with the system.

[0062] In some examples, the memory can store instances of various variable types related to the state. The variable instances can be stored with an identifier (e.g., a tag) corresponding to the type. For example, a particular identifier may be a first type "a1" for identifying a playback device of a zone, a second type "b1" for identifying a playback device that can be coupled within the zone, and a third type "c1" for identifying a zone group to which the zone can belong. As a related example, the identifier associated with the second bedroom 101c can indicate that the playback device is the only playback device in Zone C rather than within a zone group. The identifier associated with the Den can indicate that the Den is not grouped with other zones but includes the coupled playback devices 110h - 110k. The identifier associated with the dining room can indicate that the dining room is part of the dining + kitchen zone group 108b and that devices 110b and 110d are grouped (FIG. 1L). The identifier associated with the kitchen can indicate the same or similar information by virtue of the kitchen being part of the dining + kitchen zone group 108b. Other exemplary zone variables and identifiers are described below.

[0063] In yet another example, the media playback system 100 may be a variable or identifier representing other associations of zones and zone groups, such as an identifier associated with an area, as shown in FIG. 1M. The area may include clusters of zone groups and / or zones not within a zone group. For example, FIG. 1M shows an upper area 109a including zones A - D and a lower area 109b including zones E - I. In one aspect, the area may be used to call a zone group and / or a cluster of zones that share one or more zones of another cluster. In another aspect, this is different from a zone group that does not share zones with another zone group. Further examples of techniques for implementing areas can be found, for example, in U.S. Patent Application No. 15 / 682,506, filed on August 21, 2017, entitled "Association of Rooms Based on Name", and U.S. Patent No. 8,483,853, filed on September 11, 2007, entitled "Control and Operation of Grouping in a Multi - Zone Media System". Each of these applications is hereby incorporated by reference in its entirety. In some examples, the media playback system 100 may not implement an area, in which case the system may not store variables associated with the area.

[0064] III. Multi - device playback of generated media content Figure 2 is a functional block diagram of a system 200 for playing generated media content. As described above, generated media content can include any media content (e.g., audio, video, audiovisual output, tactile output, or any other media content) that is dynamically created, synthesized, and / or modified by a non-human rule-based process such as an algorithm or model. This creation or modification can be done for real-time or near real-time playback. In addition or alternatively, generated media content can be generated or modified asynchronously (e.g., before playback is requested), and specific items of generated media content can then be selected later for playback. As used herein, a “generated media module” includes any system that can generate generated media content based on one or more inputs, whether implemented in software, a physical model, or a combination thereof. In some examples, such generated media content can be created as completely new or can include new media content created by mixing, combining, manipulating, or otherwise changing one or more existing portions of media content. As used herein, a “generated media content model” includes any algorithm, schema, or set of rules that can be used to generate new generated media content using one or more inputs (e.g., sensor data, parameters provided by an artist, media segments such as audio clips or samples, etc.). In an example, the generated media module can use various different generated media content models to generate various different generated media content. In some cases, an artist or other collaborator can interact with, create, and / or update a generated media content model to generate specific generated media content.Throughout this description, several examples refer to audio content, but the principles disclosed herein can, in some examples, be applied to other types of media content, such as video, audiovisual, tactile, or others.

[0065] As shown in FIG. 2, system 200 includes a generating media group coordinator 210 that communicates with generating media group members 250a and 250b, as well as sensor data source 218, media content source 220, and control device 130. Such communication can be carried out via a network 102 that can include any suitable wired or wireless network connection or combination thereof (e.g., WiFi network, Bluetooth, Z-Wave network, ZigBee, Ethernet connection, Universal Serial Bus (USB) connection, etc.).

[0066] One or more remote computing devices 106 can also communicate with the group coordinator 210 and / or group members 250a and 250b via the network 102. In various examples, the remote computing device 106 can be a cloud-based server associated with a device manufacturer, media content provider, voice assistant service, or other suitable entity. As shown in FIG. 2, the remote computing device 106 can include a media generation module 214. As described in more detail elsewhere in this specification, the remote computing device 106 can generate media content remotely from local devices (e.g., coordinator 210 and members 250a and 250b). The generated media content can then be transmitted to one or more local devices for playback. In addition to or instead of this, the generated media content can be generated in whole or in part via local devices (e.g., group coordinator 210 and / or group members 250a and 250b). In some examples, the group coordinator 210 can itself be a remote computing device, communicatively coupled to group members 250a and 250b via a wide area network, and the devices need not be located in the same place within the same environment (e.g., home, office, etc.).

[0067] a. Example of Media Generation Group Operation In the illustrated example, the generation media group includes a generation media group coordinator 210 (also referred to herein as the "coordinator device 210") and first and second generation media group members 250a and 250b (collectively referred to herein as the "first member device 250a", the "second member device 250b", and the "member device 250"). Optionally, one or more remote computing devices 106 can also form part of the generation media group. During operation, these devices can communicate with each other and / or with other components (e.g., sensor data source 218, control device 130, media content source 220, or any other suitable data source or component) to facilitate the generation and playback of generation media content.

[0068] In various examples, some or all of the devices 210 and / or 250 can be located in the same place within the same environment (e.g., within the same home, store, etc.). In some examples, at least some of the devices 210 and / or 250 can be separated from each other, e.g., in different homes, different cities, etc.

[0069] The coordinator device 210 and / or the member device 250 can include some or all of the components of the playback device 110 or the network microphone device 120 described above with respect to FIGS. 1A - 1H. For example, the coordinator device 210 and / or the member device 250 can optionally include a playback component 212 (e.g., a transducer, amplifier, audio processing component, etc.), or such a component can be omitted in some cases.

[0070] In some examples, the coordinator device 210 is the playback device itself and can thus also operate as a member device 250. In other examples, the coordinator device 210 can be connected to one or more member devices 250 (e.g., via a direct wired connection or via the network 102), but the coordinator device 210 does not itself play the generated media content. In various examples, the coordinator device 210 can be implemented on a bridge-like device on a local network, a playback device that is not part of its own generated media group (i.e., the playback device itself does not play the generated media content), and / or a remote computing device (e.g., a cloud server).

[0071] In various examples, one or more of the devices can include a generated media module 214 thereon. Such a generated media module 214 can generate new synthetic media content based on one or more inputs, e.g., using an appropriate generated media content model. As shown in FIG. 2, in some examples, the coordinator device 210 can include a generated media module 214 for generating the generated media content, and the generated media module can then be transmitted to member devices 250a and 250b for simultaneous and / or synchronized playback. In addition to or instead of this, some or all of the member devices 250 (e.g., member device 250b shown in FIG. 2) can include a generated media module 214, and the generated media module can be used by the member device 250 to locally generate the generated media content based on one or more inputs. In various examples, the generated media content can optionally be generated via the remote computing device 106 using one or more input parameters received from a local device. This generated media content can then be transmitted to one or more local devices for adjustment and / or playback.

[0072] In some examples, at least some of the member devices 250 do not include the generation media module 214 therein. Alternatively, in some cases, each member device 250 can include the generation media module 214 therein and can be configured to locally generate generation media content. In at least some examples, none of the member devices 250 include the generation media module 214 therein. In such cases, the generation media content can be generated by the coordinator device 210. Such generated media content can then be sent to the member devices 250 for simultaneous and / or synchronized playback.

[0073] In the example shown in FIG. 2, the coordinator device 210 further includes an adjustment component 216. As will be described in more detail herein, in some cases, the coordinator device 210 can facilitate the playback of generation media content via a plurality of different playback devices (which may or may not include the coordinator device 210 itself). In operation, the adjustment component 216 is configured to facilitate synchronization of both generation media creation (e.g., using one or more generation media modules 214 that can be distributed among various devices) and generation media playback. For example, the coordinator device 210 can send timing data to the member devices 250 to facilitate synchronized playback. In addition to or instead of this, the coordinator device 210 can send inputs regarding the generation media module 214, generation media model parameters, or other data to one or more member devices 250 such that the member devices 250 can locally generate generation media (e.g., using a locally stored generation media module 214) and / or such that the member devices 250 can update or modify the generation media module 214 based on inputs received from the coordinator device 210.

[0074] As will be described in more detail elsewhere in this specification, the generation media module 214 can be configured to generate generation media based on one or more inputs using a generation media content model. The inputs can include sensor data (e.g., as provided by the sensor data source 218), user input (e.g., when received from the control device 130, or via direct user interaction with the coordinator device 210 or the member device 250), and / or the media content source 220. For example, the generation media module 214 can generate and continuously modify the generated audio by adjusting various characteristics of the generated audio based on one or more input parameters (e.g., sensor data regarding one or more users with respect to the devices 210, 250).

[0075] b. Exemplary Media Content Source The media content source 220 can include one or more local and / or remote media content sources in various examples. For example, the media content source 220 can include one or more local audio sources 105 as described above (e.g., audio received via an input / output connection from a mobile device (e.g., a smartphone, a tablet, a laptop computer) or another suitable audio component (e.g., a television, a desktop computer, an amplifier, a phonograph, a Blu-ray player, a memory storing digital media files)). In addition to, or instead of, this, the media content source 220 can include one or more remote computing devices accessible via a network interface (e.g., via communication over the network 102). Such remote computing devices can include, for example, individual computers or servers such as a media streaming service server storing audio and / or other media content.

[0076] In various examples, media available via media content source 220 can include pre-recorded audio segments in the form of full sound, songs, parts of songs (e.g., samples), or any audio component (e.g., pre-recorded audio of a particular instrument, synthetic beats or other audio segments, non-musical audio such as spoken words or natural sounds, etc.). During operation, such media can be used by generation media module 214 to generate generated media content, for example, by combining, mixing, overlaying, manipulating, or otherwise changing the retrieved media content to generate new generated media content for playback via one or more devices. In some examples, the generated media content can take the form of a combination of pre-recorded audio segments (e.g., pre-recorded songs, recordings of spoken words, etc.) and new synthetic audio created and overlaid with the pre-recorded audio. As used herein, "generated media content" or "generated media content" can include any such combination.

[0077] c. Exemplary Generation Media Module As described above, the generation media module 214 can include any system that can generate generation media content based on one or more inputs, regardless of whether it is instantiated with software, a physical model, or a combination thereof. In various examples, the generation media module 214 can utilize a generation media content model, which can include one or more algorithms or mathematical models that determine how media content is generated based on relevant input parameters. In some cases, the algorithms and / or mathematical models themselves can be updated over time based on instructions received from, for example, one or more remote computing devices (such as a cloud server associated with a music service or other entity), or inputs received from other group member devices within the same or a different environment, or any other suitable input. In some examples, various devices within the group can have different generation media modules 214 thereon, for example, using a first member device having a different generation media module 214 than a second member device. In other cases, each device within the group having the generation media module 214 can include substantially the same model or algorithm.

[0078] To generate the generation media content, any suitable algorithm or combination of algorithms can be used. Examples of such algorithms include machine learning techniques (such as adversarial generative networks, neural networks, etc.), formal grammars, Markov models, finite state automata, and / or any algorithms implemented within currently available offerings such as JukeBox by OpenAI, AWS DeepComposer by Amazon, Magenta by Google, Amper AI by Amper Music. In various examples, the generation media module 214 can utilize any suitable generation algorithm that currently exists or will be developed in the future.

[0079] In accordance with the above description, generating media content (e.g., audio content) can include changing various characteristics of the media content in real time and / or algorithmically generating new media content in real time or near real time. In the context of audio content, this can be achieved by storing several audio samples in a database (e.g., within media content source 220) that can be remotely located and accessible by coordinator device 210 and / or member devices 250 via network 102, or the audio samples can be maintained locally on devices 210, 250 themselves. The audio samples can be associated with one or more metadata tags corresponding to one or more audio characteristics of the sample. For example, a given sample can be associated with a metadata tag indicating that the sample contains audio of a particular frequency or frequency range (e.g., bass / midrange / treble) or a particular instrument, genre, tempo, key, release date, geographic region, timbre, reverb, distortion, sonic texture, or any other audio characteristic that becomes apparent.

[0080] During operation, the generation media module 214 (e.g., of the coordinator device 210 and / or the second member device 250b) can search for specific audio samples based on their associated tags and mix the audio samples to create generated audio. The generated audio can evolve in real time when the generation media module 214 searches for audio samples with different tags and / or different audio samples with the same or similar tags. The audio samples searched by the generation media module(s) 214 can depend on one or more inputs such as various user inputs like sensor data, time, geographical location, weather, or mood selection, or physiological inputs like heart rate. In this way, when the input changes, the generated audio also changes. For example, if the user selects an input for a calming or relaxing mood, the generation media module(s) 214 can search for audio samples and mix them with tags corresponding to audio content that the user can calm down or relax to. Examples of such audio samples can include audio samples tagged as having a low tempo or low harmonic complexity, or audio samples that are pre-determined and tagged as being calm and relaxing. In some examples, the audio samples can be identified as calming or relaxing based on an automated process that analyzes the temporal and spectral content of the signal. Other examples are possible as well. In any of the examples herein, the generation media module 214 can adjust the characteristics of the generated audio by searching for and mixing audio samples associated with different metadata tags or other suitable identifiers.

[0081] Changing the characteristics of the generated audio can include manipulating one or more of volume, balance, removal of specific instruments or tones, tempo of the audio, gain, reverb, spectral equalization, timbre, or sound texture. In some examples, the generated audio can be played differently on different devices, such as by emphasizing certain characteristics of the generated audio on a particular playback device closest to the user. For example, the closest playback device can emphasize a particular instrument, beat, tone, or other characteristic, while the remaining playback devices can function as background audio sources.

[0082] As described elsewhere in this specification, the media content module 214 can be configured to generate media intended to direct the user's mood and / or physiological state in a desired direction. In some examples, the user's current state (e.g., mood, emotional state, activity level, etc.) is constantly and / or repeatedly monitored or measured (e.g., at predetermined intervals) to ensure that the user's current state is transitioning towards a desired state or at least not in a direction opposite to the desired state. In such examples, the generated audio content can be varied to direct the user's current state towards a desired end state.

[0083] In any of the examples of this specification, the generation media module can use hysteresis to avoid rapid adjustments to generated audio that could negatively impact the listening experience. For example, if the generation media module changes the media based on an input of the user's position relative to the playback device, as the user approaches or moves away from the playback device, the playback device can rapidly change the generated audio in any of the ways described in this specification. Such rapid adjustments can be unpleasant for the user. To reduce these rapid adjustments, the generation media module 214 can be configured to use hysteresis by delaying adjustments to the generated audio for a predetermined period when the user's movement or other activity triggers an adjustment. For example, if the playback device detects that the user has moved within a threshold distance of the playback device, instead of immediately performing one of the aforementioned adjustments, the playback device can wait for a predetermined amount of time (e.g., several seconds) before making the adjustment. If the user remains within the threshold distance after the predetermined amount of time, the playback device can proceed with the adjustment of the generated audio. However, if the user does not remain within the threshold distance after the predetermined amount of time, the generation media module 214 can refrain from adjusting the generated audio. The generation media module 214 can similarly apply hysteresis to other generation media adjustments described in this specification.

[0084] Figure 3 shows a flowchart of a process 300 for generating audio content using various input parameters. In various examples, one or more of these input parameters can be changed based on user input. For example, an artist can select various parameters, constraints, or available audio segments shown in Figure 2, and these selections can at least partially determine the final output of the generated audio content. As described above, such a generation media module may be stored and operated on one or more playback devices for local playback (e.g., via the same playback device and / or via other playback devices communicatively coupled via a local area network). In addition or alternatively, such a generation media module may be stored and operated on one or more remote computing devices, and the resulting output is transmitted to one or more remote devices via a wide area network for playback.

[0085] As shown, the process begins at block 302 and proceeds to block 304 to a clock / metronome, where inputs for tempo 306 and time signature 308 are received. The tempo 306 and time signature 308 can be selected by an artist or automatically determined or generated using a model. The process proceeds to block 310, where a chord change can be triggered and receives a chord change frequency parameter 312 as an input. An artist can choose to have a higher chord change frequency in music intended for a higher energy experience (e.g., dance music, uplifting ambient music, etc.). Conversely, a lower chord change frequency may be associated with a lower energy output (e.g., quiet music).

[0086] In block 314, code is selected from available code segments 316. A plurality of code information parameters 318, 320, 322 can also be provided as inputs to the code segment 316. Using these inputs, the specific code to be played next can be determined and output as block 324. In some examples, an artist can provide information about each code, such as weighting, the frequency with which that specific code should be used, and so on.

[0087] Next, in block 326, a code variation is selected based at least in part on a harmonic complexity parameter that functions as an input. The harmonic complexity parameter 328 can be adjusted or selected by an artist, or can be determined automatically. Generally, a higher harmonic complexity parameter may be associated with a higher energy audio output, and a lower harmonic complexity parameter may be associated with a lower energy audio output. In some cases, the harmonic complexity parameter can include inputs such as code inversion, articulation, and harmonic density.

[0088] In block 330, the process obtains the root of the code and, in block 332, selects the bass segment to be played from among the available bass segments 334. These bass segments then undergo bus processing 336, where equalization, filtering, timing, and other processing can be performed.

[0089] Returning to the code variation in block 326, the process then continues separately in block 338 to play the harmony selected from among the available harmony segments 340. This harmony segment then undergoes bus processing 342. The harmony segment bus processing 342 can perform processing such as equalization, filtering, timing, etc., similar to the bass bus processing.

[0090] When returning to the selected code 324, the process continues to individually filter melody notes at block 344 using the input of melody constraint 346. The output at block 348 is the melody notes available for playback. Melody constraint 346 can be provided by the artist and can specify, for example, which sounds to play, limit the melody range, or provide other such constraints that may depend on the specific selected code 324.

[0091] At block 350, the process determines which melody notes to play (from among the available melody notes 348). This determination can be made automatically based on model values, artist-provided inputs, randomization effects, or any other appropriate inputs. In the illustrated example, one input is from the trigger melody note block 352, which is based on the melody density parameter 354. The artist can provide the melody density parameter 354, which partially determines how complex and / or energetic the audio output is. Based on that parameter, the block 352 input to block 350 is used to determine which melody notes to play, and melody notes may be triggered more or less frequently and at specific times. In various examples, the output of block 350 may be provided as an input to block 350 in the form of a feedback loop such that the next melody note selected at block 350 depends at least in part on the melody note last selected at block 350. Next, at block 356, a melody segment is selected from among the available melody segments 358, and then the melody segment undergoes bus processing 360.

[0092] Returning to the start of block 302, the process proceeds separately to block 362 to play non-music content. This may be, for example, natural sounds, spoken word audio, or other such non-music content. Various non-music segments 364 can be stored and played. These non-music content segments can also undergo bus processing at block 366.

[0093] The outputs of these various paths (e.g., selected bass segment(s), harmony segment(s), melody segment(s), and / or non-music segment(s)) can each undergo separate bus processing before being combined at block 368 via mixing and mastering processing. Here, it is possible to set the combined level, apply various filters, establish relative timing, and perform any other appropriate processing steps that are carried out before the generated audio content is output at block 370. In various examples, some of the paths may be completely omitted. For example, the generation media module may omit the option to play non-music content along with the generated music content. The process 300 shown in FIG. 3 is merely exemplary, and as will be understood by those skilled in the art, appropriate changes can be made to the process 300 shown here, and furthermore, there are numerous appropriate alternative processes that can be used to generate media content.

[0094] FIG. 4 is an exemplary architecture for storing and retrieving generated media content. In this example, the generated media content includes various individual tracks (each having multiple variations related to energy level or another parameter) that can be selected and played in various orders and groupings according to specific input parameters.

[0095] As shown, the generated media content 404 can be stored as one or more audio files associated with the global generated media content metadata 402. Such metadata can include, for example, a global tempo (e.g., beats per minute), a global trigger frequency (e.g., how often to check for changes in input parameters), and / or a global crossfade duration (e.g., the time to fade between different selected energies).

[0096] Among the generated media content 404, there are a plurality of different tracks 406, 408, and 410. During operation, these tracks can be selected and played in various arrangements (e.g., randomized grouping by some overlay, or playing according to a predetermined sequence, etc.). In some examples, the generated media content 404 including the tracks 406, 408, 410 can be stored locally via one or more playback devices, while one or more remote computing devices can periodically send updated versions of the tracks, the generated media content, and / or the global generated media content metadata. In some examples, the remote computing device can be periodically polled or queried by the playback device, and in response to the query or poll, the remote computing device can supply updates to the generated media module stored on the local playback device.

[0097] For each track, there can be corresponding subsets of that track corresponding to different energy levels. For example, the first energy level (EL) of track 1 at 412, the second energy level of track 1 at 414, and the n energy level of track 1 at 416. Each of these can include both metadata corresponding to a particular energy level (e.g., metadata 418, 420, 422) and particular media files (e.g., media files 424, 426, 428). In some examples, each track can include a plurality of media files (e.g., media file 424) arranged in a particular way, and the arrangement and combination can be desired by the corresponding metadata (e.g., metadata 418). The media files can be, for example, in any suitable format that can be played via a playback device and / or streamed to a playback device for playback. In some examples, one or more of media files 424, 426, 428 can be the output of the generation model shown in FIG. 3. The metadata can include, for example, tempo (if different from the global tempo), trigger frequency (if different from the global trigger frequency), sequence information (e.g., whether to play specific files in order, randomly, or with percent weighting), crossfade duration (if different from the global crossfade), spatial information (e.g., for rendering audio content in space using multiple transducers), polyphony information (e.g., enabling multiple audio files to be played at once in this segment), and / or level (e.g., level adjustment in dB units, or random within a predetermined range).

[0098] During operation, one or more input parameters (e.g., the number of people present in the room, the time, etc.) can be used to determine a target energy level. This determination can be made using the playback device and / or one or more remote computing devices. Based on this determination, a specific media file corresponding to the determined energy level can be selected. Next, the generating media module can arrange and play those selected tracks according to a generating content model. This can include playing the selected tracks in a specific predetermined order, in a random or pseudo-random order, or any other suitable method. The tracks can be played in a way that at least partially overlaps in some examples. It may be useful to vary the amount of overlap between tracks so that a casual listener does not hear a repeating loop of audio content but instead perceives the generated audio as an infinite stream of non-repeating audio.

[0099] The example shown in FIG. 4 utilizes the energy level as a parameter for differentiating different generated audio content, but in various examples, specific variations or permutations of the generated audio content can vary along other dimensions (e.g., genre, time, associated user tasks, etc.).

[0100] d. Exemplary sensor data sources and other input parameters As described above, the generation media module 214 can generate generation media based at least in part on input parameters that can include sensor data (e.g., when received from the sensor data source 218) and / or other suitable input parameters. With respect to sensor input parameters, the sensor data source 218 can include data from any suitable sensor, regardless of where it is located with respect to the generation media group, and regardless of the value measured thereby. Examples of suitable sensor data include physiological sensor data such as data obtained from biometric sensors, wearable sensors, and the like. Such data can include physiological parameters such as heart rate, respiratory rate, blood pressure, brain waves, activity level, movement, body temperature, and the like.

[0101] Suitable sensors include wearable sensors configured to be worn or carried by a user, such as a headset, a watch, a mobile device, a brain machine interface (e.g., Neuralink), headphones, a microphone, or other similar devices. In some examples, the sensor can be a non-wearable sensor or can be fixed to a stationary structure. The sensor can provide sensor data that can include data corresponding to, for example, brain activity, voice, position, movement, heart rate, pulse, body temperature, and / or sweating. In some examples, the sensor can correspond to multiple sensors. For example, as described elsewhere in this specification, the sensor can correspond to a first sensor worn by a first user, a second sensor worn by a second user, and a third sensor not worn by the user (e.g., fixed to a stationary object or structure). In such an example, the sensor data can correspond to a plurality of signals received from each of the first, second, and third sensors.

[0102] The sensor can be configured to obtain or generate information that generally corresponds to the user's mood or emotional state. In one example, the sensor is a wearable brain sensing headband, which is one of many examples of sensors described herein. Such a headband can include, for example, an electroencephalogram (EEG) headband having a plurality of sensors thereon. In some examples, the headband can correspond to either the Muse (trademark) headband (InteraXon; Toronto, Canada). The sensors can be placed at various positions around the inner surface of the headband so as to correspond to different anatomical structures of the user's brain (e.g., frontal bone, parietal bone, temporal bone, and sphenoid bone). In this way, each sensor can receive different data from the user. Each of the sensors can correspond to an individual channel that can be streamed from the headband to system devices 210 and / or 250. Such sensor data can be used to detect the user's mood, for example, by classifying the frequencies and intensities of various brain waves or by performing other analyses. Further details regarding the use of the brain sensing headband for generating audio content can be found in co-owned U.S. Patent Application No. 62 / 706,544, filed on August 24, 2020, entitled MOOD DETECTION AND / OR INFLUENCE VIA AUDIO PLAYBACK DEVICES, which is incorporated herein by reference in its entirety.

[0103] In some examples, the sensor data source 218 includes data obtained from sensors of network devices (e.g., Internet of Things (IoT) sensors such as networked lights, cameras, temperature sensors, thermostats, presence detectors, microphones, etc.). In addition to or instead of this, the sensor data source 218 can include environmental sensors (e.g., those that measure or display weather, temperature, time / day / week / month, etc.).

[0104] In some examples, the generative media module 214 can utilize inputs in the form of playback device capabilities (e.g., number and type of transducers, output power, other system architectures), device location (e.g., location relative to one or more users or relative to other playback devices). Further examples of generating and modifying generated audio as a result of user and device location are described in detail in U.S. Patent Application No. 62 / 956,771, filed on January 3, 2020, titled "GENERATIVE MUSIC BASED ON USER LOCATION", the entire disclosure of which is incorporated herein by reference. Additional inputs can include device states of one or more devices in the group, such as thermal state (e.g., if a particular device is at risk of overheating, the generated content can be modified to lower the temperature), battery level (e.g., bass output can be reduced in a portable playback device with a low battery level), and coupling state (e.g., whether a particular playback device is configured as part of a stereo pair, coupled to a subwoofer, configured as part of a home theater system, etc.). Any other suitable device characteristics or states can similarly be used as inputs for generating the generative media content.

[0105] Another exemplary input parameter includes the presence of users. For example, when a new user enters a space where generated audio is being played, the presence of the user can be detected (e.g., via proximity sensors, beacons, etc.), and the generated audio can be modified in response. This modification can be based on the number of users (e.g., ambient meditative audio for one user, relaxing music for two to four users, and party or dance music for more than four users is presented). The modification can also be based on the identification information of the users present (e.g., a user profile based on the characteristics of the users, listening history, or other such credentials).

[0106] In one example, the user can wear a biometric device that can measure various biometric parameters such as the user's heart rate or blood pressure and report those parameters to devices 210 and / or 250. The generation media module 214 of these devices 210 and / or 250 can use these parameters to further adapt the generated audio, for example, by increasing the tempo of the music in response to the detection of a high heart rate (since it can indicate that the user is engaged in high physical activity), or by decreasing the tempo of the music in response to the detection of high blood pressure (since it can indicate that the user is stressed and can benefit from calming the music).

[0107] In yet another example, one or more microphones of a playback device (e.g., microphone 115 of FIG. 1F) can detect the user's voice. The captured voice data can then be processed to determine, for example, the user's mood, age, or gender, identify a particular user from among multiple users in a household, or identify any other such input parameter. Other examples are similarly possible.

[0108] e. Example of adjustment among group members Figure 5 is a functional block diagram showing data exchange in a system for playing generated media content. For purposes of explanation, the system 500 shown in FIG. 5 includes the interaction between the coordinator device 210 and the member device 250b. However, the interactions and processes described herein can be applied to interactions involving multiple additional coordinator devices 210 and / or member devices 250. As shown in FIG. 5, the coordinator device 210 includes a generated media module 214a that receives inputs including input parameters 502 (e.g., sensor data, media content, model parameters of the generated media module 214a, or other such inputs) and clock and / or timing data 504. In various examples, the clock and / or timing data 504 can include synchronization signals for synchronizing playback and / or for synchronizing generated media generated by various devices within a group. In some examples, the clock and / or timing data 504 can be provided by an internal clock, a processor, or other such components housed within the coordinator device 210 itself. In some examples, the clock and / or timing data 504 can be received from a remote computing device via a network interface.

[0109] Based on these inputs, the generated media module 214a can output generated media content 404a. Optionally, the output generated media content 404a itself can function as an input to the generated media module 214a in the form of a feedback loop. For example, the generated media module 214a can use a model or algorithm that depends at least in part on previously generated content to generate subsequent content (e.g., audio frames).

[0110] In the illustrated example, the member device 250b may similarly include a generation media module 214b that may be substantially the same as the generation media module 214a of the coordinator device 210, or may be different in one or more aspects. The generation media module 214b can similarly receive input parameters 502 and clock and / or timing data 504. These inputs can be received from the coordinator device 210, from other member devices, from other devices on the local network (e.g., a locally network-connected smart thermostat that supplies temperature data), and / or from one or more remote computing devices (e.g., a cloud server that provides clock and / or timing data 504, or weather data, or any other such input). Based on these inputs, the generation media module 214b can output generated media content 404b. This generated generated media content 404b can optionally be fed back to the generation media module 214b as part of a feedback loop. In some examples, the generated media content 404b can include or consist of the generated media content 404a transmitted to the member device 250b via the network (generated via the coordinator device 210). In other cases, the generated media content 404b may be generated separately and independently from the generated media content 404a generated via the coordinator device 210.

[0111] Next, the generated media contents 404a and 404b can be played via the devices 210 and 250b themselves and / or by other devices within the group. In various examples, the generated media contents 404a and 404b can be configured to be played simultaneously and / or synchronously. In some cases, the generated media contents 404a and 404b may be substantially identical or similar to each other, and each generated media module 214 utilizes the same or similar algorithms and the same or similar inputs. In other examples, the generated media contents 404a and 404b are still configured for synchronous or simultaneous playback but may be different from each other.

[0112] f. Exemplary Generated Media Using a Distributed Architecture As described above, the generation of media content can be computationally intensive and, in some cases, it may be unrealistic to perform entirely on a local playback device alone. In some examples, the generation media module of the local playback device can request generation media content from a generation media module stored on one or more remote computing devices (e.g., a cloud server). The request can include or be based on specific input parameters (e.g., sensor data, user input, context information, etc.). In response to the request, the remote generation media module can stream specific generation media content to the local device for playback. The specific generation media content provided to the local playback device can change over time depending on specific input parameters, the configuration of the generation media module, or other such parameters. In addition to or instead of this, the playback device can store individual tracks for playback (e.g., as shown in FIG. 4, different variations of the track are associated with different energy levels). The remote computing device can then periodically provide new files for the updated tracks to the local playback device for playback, or can provide updates to the generation media module that determines how and when to play specific files stored locally on the playback device.

[0113] In this way, the tasks necessary to generate and play the generated audio are distributed between one or more remote computing devices (s) and one or more local playback devices. By performing at least some of the computationally intensive tasks associated with the generation of new media content for the remote computing device, and optionally by reducing the need for real-time calculations, the overall efficiency can be improved. By generating a discrete number of alternative tracks or track variations according to a particular media content model prior to playback via the remote computing device, the local playback device can request and receive particular variations based on real-time or near real-time input parameters (e.g., sensor data). For example, the remote computing device can generate different versions of the media content, and the playback device can request a particular version in real time based on the input parameters. As a result, appropriate generated media content is played back based on real-time or near real-time input parameters (e.g., sensor data) without the need for de novo generation of such media content to be performed in real time.

[0114] FIG. 6 is a schematic diagram of an exemplary distributed generation media playback system 600. As shown, artist 602 can supply a plurality of media segments 604 and one or more generation content models 606 to a generation media module 214 stored via one or more remote computing devices. The media segments can correspond to, for example, specific audio segments or seeds (e.g., individual notes or chords, short tracks of n measures, non-musical content, etc.). In some examples, the generation content model 606 can also be supplied by artist 602. This can include providing the entire model, or artist 602 can provide an input to model 606, for example, by changing or adjusting specific aspects (e.g., tempo, melody constraints, harmonic complexity parameters, chord change density parameters, etc.).

[0115] The generation media module 214 can receive both media segments 604 and one or more input parameters 502 (as described elsewhere in this specification). Based on these inputs, the generation media module 214 can output generation media. As shown in FIG. 6, artist 602 can optionally audition the generation media module 214 by receiving an exemplary output based on, for example, an input provided by artist 602 (e.g., media segment 604 and / or generation content model 606). In some cases, the audition can play back variations of the generation media content to artist 602 in response to various different input parameters (e.g., one version corresponds to a high energy level intended to produce an exciting or uplifting effect, and another version corresponds to a low energy level intended to produce a calming effect, etc.). Based on the output from this audition step, artist 602 can dynamically update the settings of media segment 604 and / or generation content model 606 until the desired output is achieved.

[0116] In the illustrated example, there can be an iteration at block 608 every n hours (or minutes, days, etc.) where the generation media module 214 can generate multiple different versions of the generated media content. In the illustrated example, there are three versions: version A of block 610, version B of block 612, and version C of block 614. These outputs are stored as generated media content 616 (e.g., via a remote computing device). A particular version (version C as block 618 in this example) can be sent (e.g., streamed) to the local playback device 250 for playback. In some examples, a particular version can correspond to tracks 406, 408, and 410 shown in FIG. 4.

[0117] Although three versions are shown here as an example, in reality, there may be more versions of the generated media content generated via a remote computing device. The versions can vary along several different dimensions, such as being suitable for different energy levels, being suitable for different intended tasks or activities (e.g., study vs. dance), being suitable for different times, or any other appropriate variations.

[0118] In the illustrated example, the playback device 250 can periodically request a specific version of the generated media content from the remote computing device. Such requests can be based on, for example, user input (e.g., user selection via a controller device), sensor data (e.g., the number of people present in a room, background noise level, etc.), or other suitable input parameters. As shown, the input parameter 502 can optionally be provided (or detected) to the playback device 250. In addition to or instead of this, the input parameter 502 can be provided (or detected) to the remote computing device 106. In some examples, the playback device 250 sends the input parameter to the remote computing device 106, and the remote computing device provides the appropriate version to the playback device 250 without the playback device 250 specifically requesting a particular version. g. Example of a method for generating digital content based on blockchain data

[0119] As described above, a system for generating and playing media content may interact with blockchain data (or data stored via other distributed ledger technologies). For example, as shown in FIG. 6, the blockchain layer 620 can be utilized to provide data as an input to other components of the media generation and playback system 600, such as the playback device 250, input parameter 502, media segment 604, generated content model 606, and / or generation media module 214. In various embodiments, the blockchain layer 620 can store data that can be used as one or more input parameters 502, data included or usable to obtain or affect a particular media segment 604, data included or usable to obtain or affect a particular generated content model 606, and / or data included or usable to obtain or affect a particular generation media module 214. Further, some or all of these components can communicate with the blockchain layer 620 to write data to the blockchain, record a transaction, or even interact with the blockchain layer 620. For example, the playback device 250 can record a transaction reflecting the playback of a particular track on the blockchain layer 620. The data stored via the blockchain layer 620 can include a particular input parameter 502, or appropriate input parameters 502 can be generated using the data stored via the blockchain layer 620. Similarly, a particular media segment 604, generated content model 606, generation media module 214, input parameter 502, or other appropriate data can be written to the blockchain layer 620 to create an immutable record of such content, transaction, or other data. Additional details regarding the utilization of blockchain technology (or other appropriate distributed ledger technology) in the creation and playback of generated media content will be described in more detail below.

[0120] FIG. 7 is a schematic diagram of a distributed generation media playback system 700 of another example. As shown, system 700 obtains, generates, or stores, for example, input parameter 502, generation content model 606, or other data or parameters used in the creation and playback of generated media content by including or communicating with blockchain layer 620.

[0121] Examples of such a blockchain layer 620 include public distributed ledgers such as Ethereum, Bitcoin, Solana, Avalanche, Polygon, etc. Although blockchain layer 620 is shown, in various embodiments, any suitable distributed ledger technology can be used, including private or semi-private blockchains, as well as non-blockchain implementations such as directed acyclic graphs (DAGs) (e.g., Nano, IOTA, etc.). In various examples, participants using blockchain layer 620 can transact with each other in a peer-to-peer manner, and the operation of blockchain layer 620 can be decentralized so that no single central entity controls the operation of the network. Such distributed ledgers can be used to track the creation, exchange, and repayment of specific real-world assets such as currency. This approach enables strong auditing of asset transactions due to the practical immutability of the data stored on the blockchain. Currency is just one of the various assets that are desirable to track with a distributed ledger. Other types of assets may differ from currency with respect to one or more operations that manage the creation, exchange, and / or repayment of the asset. Additionally, different blockchain architectures can also differ with respect to policies, protocols, and even the tools used to program the behavior of the asset.

[0122] Generally, a distributed ledger in which each unit of an asset is represented by some form of digital token can be programmed to confer a set of behaviors appropriate to the asset that the token represents. By way of example, "fungible" behavior enables an asset to be exchanged for other assets of the same class. A unit of currency of a certain denomination (e.g., one dollar) is fungible because all units have the same value as other units of the same denomination. In contrast, a property title is "non-fungible" because its value depends on the size, location, and other aspects of the specified property. For each asset represented as a token, appropriate fungible or non-fungible behavior is programmed into the class of tokens of the virtual ledger that tracks the asset.

[0123] In some embodiments, the tokens traded via the blockchain layer 620 are non-fungible. Such non-fungible tokens (NFTs) are unique and non-exchangeable with other tokens. NFTs can consist of and / or be associated with unique digital artworks and / or music, domain names, digital collectibles (such as CryptoKitties, memes, etc.), event tickets, parts of virtual worlds, digital objects used in games, avatars or characters, items with utility (such as providing voting rights or governance rights), and the like. In various examples, an NFT may itself contain associated data (e.g., the raw audio data of a music NFT may be stored on-chain in some cases), or it may contain a pointer (such as a URL or URI) that directs to data stored elsewhere (e.g., audio data stored on a server managed by the issuer of a music NFT).

[0124] Such tokens, whether fungible or non-fungible, can be stored by the user via a digital wallet. A digital wallet is a device, physical medium, program, or service that can store the public key and / or private key of a blockchain transaction. In some examples, a digital wallet can store multiple public key / private key pairs for various different blockchains, and a user can store assets related to different blockchains in a single wallet. Examples include MetaMask, Phantom, Coinbase Wallet, Ledger Nano, etc. In operation, the user can sign blockchain transactions via the wallet using the appropriate private key (or by permitting the wallet to sign transactions with the private key). Thereafter, if the transaction signature is valid, the transaction is confirmed and added to the corresponding block of the blockchain. In some examples, the wallet identifier itself may be used as an input parameter to a generative model, independent of the tokens held in a particular wallet.

[0125] In various implementations, the blockchain layer 620 can be configured to automatically execute transactions under one or more conditions. Such self-executing transactions can be referred to as "smart contracts." A smart contract is computer code stored on a blockchain and is configured to be executed only under specific circumstances or in a specific manner. For example, a smart contract can be configured such that at a certain time, based on one or more other transactions or other appropriate criteria, a specific transaction is executed when a certain threshold is exceeded. In some examples, the generation media module 214 and / or the generation content model 606 can be implemented in the form of a smart contract, whereby by interacting with the smart contract via the blockchain layer 620, the smart contract can output generated media content, a generated content model, or data or instructions that can be used to generate such generated media content or generate a generated content model.

[0126] One organizational structure unique to blockchains is the decentralized autonomous organization (DAO). A DAO is typically a community-driven entity without a central authority. Such a DAO becomes fully autonomous and transparent as smart contracts provide the basic rules and execute agreed-upon decisions. Community voting can be conducted by token holders using on-chain transactions. Based on specific voting results, the smart contract can execute specific transactions or other code to implement the decisions of DAO members. Generally, a DAO issues tokens to users in exchange for currency investments or donations, or for free (e.g., via an "airdrop"). Token holders usually retain a certain voting right, which may be proportional to the holding amount. In some cases, token holders also receive financial benefits such as a share of the transaction fees collected by the DAO.

[0127] In the exemplary system 700 shown in FIG. 7, the generative media module 214 can receive a number of different inputs and output one or more generative content versions 610 in response. These content versions 610 are stored in the generative media content store 616, from which a particular selected generative content version 618 can be selected and played via the playback device 250 or other output device (e.g., the optical component of the generative media content version 618 can cause the lighting device 702 to output light, in which case a particular hue, color temperature, brightness, on / off or other pattern, etc., follows the generative media content version 618). The inputs to the generative media module 214 include a generative content model 606 and input parameters 502, similar to the approach described above with respect to FIG. 6. As described above, in some embodiments, the generative content model 606 can be stored via the blockchain layer 620 or obtained from data stored via the blockchain layer 620. Similarly, one or more input parameters 502 include or are based on data stored via the blockchain layer 620. For example, the data of the blockchain can be used in the generative media module 214 to generate an appropriate output such as "sonification" of a data stream (e.g., a real-time feed such as the price value of a cryptocurrency can be made into a corresponding sound output). In some examples, the data of the blockchain can include data provided by one or more "oracles", which are typical third-party services that provide external information (e.g., price feeds, weather data, election results, etc.) to smart contracts. In some examples, the generative media module 606 and / or the generative media content 616 may be stored locally via the playback device 250, in which case the input parameters 502 are streamed to the playback device 250 and used to generate a new version 610 of the generative media content 616.

[0128] Additionally or alternatively, the generative media module 214 can receive as input the output of one or more smart contracts 706. In some cases, the generative media module 214 can itself take the form of a smart contract, in which case the program code is stored on the blockchain and automatically executed under certain conditions (e.g., the user 708 interacts with the smart contract 706 and, in response, certain generative media content is sent to a specified destination). In some examples, the generative media module 214 can be executed locally (rather than as a smart contract on the blockchain) or via a remote server, but can communicate with the smart contract 706 to receive input parameters 502 from the smart contract 706 or provide an appropriate output to the smart contract 706. For example, the specific generative media content output by the generative media module 214 can be used to generate one or more NFTs via the smart contract 706. In some embodiments, each specific version of the generated content generated by the generative media module 214 can take a corresponding NFT, thereby making each NFT generated by the smart contract 706 based on the input from the generative media module 214 unique. This is shown in FIG. 7, where a plurality of discrete NFTs 710a-f, indicated as NFT1 to NFTn, are generated via the smart contract 706. These NFTs 710 can further be provided as input to the generative media module 214. For example, the generative media module 214 can generate dynamically different content based at least in part on the specific NFTs 710 that interact. Additionally or alternatively, the user 708 can access the specific generative media module 214 only if the user holds the appropriate NFT 710 in a digital wallet.

[0129] In some examples, one or more NFTs 710 are existing NFTs each owned by a third - party person or entity (e.g., a person or entity not associated with user 708) and are temporarily accessible by the generative media module 214. In certain examples, the one or more NFTs 710 are not composed of audio data, but instead are composed of another data type (e.g., video, image, or other data), and these data are "sonified" or converted into a form in which the generative media module 214 (or another suitable component) can generate media content.

[0130] In the illustrated example, the smart contract 706 can also output an NFT 712 represented by the NFT0 held by the user 708. Further, this NFT 712 can interact with or be generated through the artist DAO 714, and this artist DAO 714 can communicate with one or more smart contracts 706. As described above, a DAO is generally a community - led organization where members hold tokens (e.g., NFT 712), membership is specified by the tokens, voting rights and other governance rights are provided, and additionally economic benefits such as future DAO revenues are given to the token holders. In some examples, the artist DAO 714 can distribute a portion of royalties (music usage fees) to the holders of the appropriate NFT or other tokens when they are received (e.g., user 708 can receive economic benefits from the artist DAO 714 based at least in part on the ownership of the user's NFT 712).

[0131] As an option, data corresponding to the NFT712 can be stored or embedded via a physical medium. For example, as shown in FIG. 7, the data corresponding to the NFT712 can be embedded in the record disc 716 (e.g., via a unique QR code (registered trademark), a code embedded in the groove of the record disc 716), or the data can be stored using any other suitable technology via another physical medium (e.g., NFC or other RF tag). Although the record disc 716 is illustrated, in various embodiments as other examples, the physical substrate can take various forms such as, for example, a playback device, a physical card or ticket, a poster, etc.

[0132] By using the blockchain layer 620, smart contracts 706, DAO 714, and / or NFTs 710 and 712, several advantages can be provided to the generative media playback system 700. For example, by associating a specific NFT with generative media content (e.g., a soundscape), the user 708 can obtain a personalized history and can also transfer it via the decentralized peer-to-peer trading mechanism of the blockchain layer 620. However, a problem with this approach is that the data contained in the NFT is generally static and is contrary to the dynamic data of the generative soundscape and other generative media content. Another problem related to the NFT is broken links, where the locator in the NFT no longer references the artwork associated with the NFT and the data in the NFT becomes from a previous generation. One way to reduce the possibility of broken links is to store the generative media content or the generative media engine in the blockchain layer 620. Additionally, the artwork itself may be embedded in the NFT (e.g., the artwork data is stored on-chain rather than on a separate server).

[0133] In some cases, the NFT 710 can include a specific seed used by the generation media module 214. Examples of such seeds include the media segment 604 of FIG. 6, the track of FIG. 4, the energy level, or the metadata, or any of the various components shown in FIG. 3. Optionally, the characteristics of a specific NFT (at least with respect to its use by the generation media module 14) may depend on its transaction history. For example, depending on when a particular NFT was last traded and how many times it was traded, the specific seed associated with that NFT may change dynamically. Additionally, or alternatively, different combinations of NFTs 710 connected to the generation media module 214 can result in different generated content versions 610, and a specific generation media output depends on which of the NFTs 710 are used as input.

[0134] In some examples, additional data related to the generation media playback system 700 can be stored via the blockchain layer 620. For example, the listening history of user 708 can be saved in the blockchain layer 620 to provide an immutable record of the listening history. This can include the viewing history of generation media content or non-generation content (such as standard recorded audio tracks and other content). In some cases, the "followers" of a particular user 708 can subscribe to that user's listening history by accessing the data stored via the blockchain layer 620. Since blockchains are generally permissionless and highly transparent, followers can freely access the listening history (or other content data associated with a particular network address). In yet another example, followers can subscribe to the generation media content of a particular user 708. Thus, while the generation media content of user 708 is dynamically created based on various inputs, this same media content can be enjoyed by other followers.

[0135] For example, consider an artist who wants to create a specific soundscape using the generation media module 214. The artist's fans can listen to the soundscape in real-time or near real-time via the blockchain data. In some cases, followers can use their own local generation media module 214 (or one running on another device) to generate corresponding generation media content using inputs, pointers, or other data. In at least some examples, such a local generation media module 214 can also utilize additional local inputs (such as specific playback device characteristics, local sensor data, etc.). As a result, the artist's generation media content is merged with that generated by the user's own local generation media module 214, and generation media content that is somewhat modified while reflecting the artist's intent is generated. Optionally, the artist's followers may be required to hold a specific NFT or other token in order to access the artist's generation media content. In yet another example, only users holding specific NFTs or tokens can access certain playlists or radio stations.

[0136] h. Exemplary Methods for Generation and Playback of Generated Audio Figures 8-13 are flow diagrams of exemplary methods for playing generated audio content via a plurality of separate playback devices. Methods 800, 900, 1000, 1100, 1200, 1300 can be implemented by any of the devices or systems described herein, or any other device or system currently known or later developed.

[0137] Various examples of methods 800, 900, 1000, 1100, 1200, 1300 include one or more operations, functions, or actions indicated by blocks. Although the blocks are shown in a sequential order, these blocks may also be executed in parallel and / or in an order different from the order disclosed and described herein. Also, various blocks may be combined into fewer blocks, divided into additional blocks, and / or removed based on the desired implementation.

[0138] Furthermore, for methods 800, 900, 1000, 1100, 1200, 1300, and other processes and methods disclosed herein, the flowcharts illustrate the functions and operations of possible implementations of some examples. In this regard, each block may correspond to a portion of a module, segment, or program code that includes one or more instructions executable by one or more processors to perform a particular logical function or step in the process. The program code may be stored in any type of computer-readable medium, such as a storage device including, for example, a disk or a hard drive. The computer-readable medium can include non-transitory computer-readable media such as tangible non-transitory computer-readable media that store short-term data, such as register memory, processor cache, and random access memory (RAM). The computer-readable medium can also include non-transitory media such as secondary or persistent long-term storage devices, such as read-only memory (ROM), optical or magnetic disks, and compact disc read-only memory (CD-ROM). Also, the computer-readable medium may be any other volatile or non-volatile memory system. The computer-readable medium may be regarded as, for example, a computer-readable storage medium or a tangible storage device. Furthermore, for the methods and other processes and methods disclosed herein, each block in FIGS. 8-13 may correspond to a circuit wired to perform a particular logical function within the process.

[0139] Referring to FIG. 8, method 800 begins at block 802, which includes receiving a command to play generated media content via a group or bonded zone of playback devices. Such a command can be received, for example, via control device 130 or other suitable user input.

[0140] At block 804, method 800 includes the group coordinator device providing timing information to the generating group member devices. The timing information can include context timing data (e.g., time data associated with sensor input or other user input), generated media playback timing data (e.g., timestamps and synchronization data to facilitate synchronous playback of generated media), and / or media content stream timing data based on a common clock.

[0141] At block 806, the method optionally includes determining a generated media content model used to generate the generated media. Such a model can be implemented, for example, in media content module 214 described above with respect to FIGS. 2-6. In some examples, each of the member devices can utilize the same or substantially the same generated media content model, and in other cases, some or all of the member devices can utilize different generated media content models from each other. For example, a first generated media content model can generate rhythmic beats, and a second generated media content model can generate ambient natural sounds. When played simultaneously, the generated audio generated by these different generated media content models can create a pleasant listening experience for the user. In some examples, the selection of a particular generated media content model itself can be based on one or more input parameters such as device capabilities, device location, number of users present, user sensor data, etc.

[0142] In block 808, method 800 includes the coordinator device and the member devices receiving context and / or other input data. For example, the input data can include sensor data, user input, context data, or any other relevant data that can be used as input for generating a media content model.

[0143] In block 810, method 800 continues with the coordinator device and the member devices synchronously generating and playing back generated media content.

[0144] FIG. 9 shows another method 900 for playing back generated audio content via a plurality of playback devices. Method 900 begins in block 902 with the group coordinator device receiving one or more input parameters. As described above, the input parameters can include sensor data, user input, context data, or any other input that can be used by the generation media module to generate the generated audio for playback.

[0145] In block 904, the coordinator device transmits the input parameters to one or more discrete playback devices having a generation media module. For example, the coordinator device can obtain sensor data and other input parameters and transmit them to a plurality of individual playback devices within the environment or to a plurality of individual playback devices distributed across a plurality of environments. In some examples, these input parameters can include characteristics of the generated content model itself that provide instructions, for example, to update the generation media module stored locally by one or more of the individual playback devices.

[0146] In block 906, the method includes transmitting timing data from a coordinator device to a playback device. The timing data can include, for example, clock data or other synchronization signals configured to facilitate adjustment of the generation of the generated media content and synchronous playback of the generated media content via separate playback devices.

[0147] In block 908, method 900 continues to simultaneously play the generated media content via the playback devices, at least partially based on the input parameters. As described above, the various playback devices may play the same generated audio, or each may play separate generated audio that produces a desired psychoacoustic effect for the users present when played synchronously.

[0148] In the example of FIG. 9, the generated media content can be locally generated by discrete playback devices, each of which generates and plays its own generated audio content in parallel with each other. In another method 1000 shown in FIG. 10, the generated media content is generated by a coordinator device, and then the coordinator device transmits the generated media content along with timing data to separate playback devices for synchronous playback.

[0149] In block 1002, method 1000 includes receiving one or more input parameters at a group coordinator device. Examples of input parameters are described elsewhere in this specification and include sensor data, user input, context data, or any other input that can be used by a generated media module to generate generated audio for playback.

[0150] In block 1004, the coordinator device generates first and second generated media streams based at least in part on input parameters, and in block 1006, the first and second media streams are each sent to first and second separate playback devices. For example, the coordinator device can generate two streams that form different channels of generated audio, such as a left channel played by a first playback device and a corresponding right channel played by a second playback device. In addition to or instead of this, the two streams can be separate audio tracks that can be played synchronously despite, for example, a rhythmic beat in one stream and ambient natural sounds in the other stream. Multiple other variations are possible. This example describes two streams for two playback devices, but in various other examples, there may be one stream or three or more streams that can be provided to any number of playback devices for synchronous playback. In at least some examples, one or more of the playback devices can be located in different environments that are far apart from each other (e.g., different homes, different cities, etc.).

[0151] In block 1008, the first playback device plays the first generated media stream and the second playback device currently plays the second generated media stream. In some examples, this simultaneous playback can be facilitated by using timing data received from the coordinator device.

[0152] Figure 11 shows another exemplary method 1100 for generating and playing generated media content. As described above, it may be beneficial to use one or more remote computing devices (e.g., cloud-based servers) to perform at least a portion of the processing required to generate the generated media content, reducing the computational requirements imposed on the local playback device and / or performing operations that are not possible using the components of the local playback device. Method 1100 begins, in block 1102, at the playback device, by receiving one or more input parameters. As described above, the input parameters can include sensor data, user input, context data, or any other input that can be used by the generated media module to generate generated audio for playback.

[0153] In block 1104, method 1100 includes accessing a library that includes a plurality of existing media segments. For example, a plurality of individual media segments (e.g., audio tracks) can be stored on the playback device and arranged and / or mixed for playback according to a generated content model. In addition to or instead of this, the library can be stored on one or more remote computing devices, and the individual media segments are sent from the remote computing device to the playback device for playback.

[0154] Method 1100 continues, at block 1106, to generate media content based at least in part on input parameters by placing a selection of existing media segments from a library for playback according to a generated media content model. As described elsewhere herein, the generated media content model can receive one or more input parameters as input. Based on the input, the generated media content model can be used to output a particular generated media content. In an example, the generated media content can include the placement of existing media segments, for example, arranging them in a particular order, with or without duplication between particular media segments, and / or with additional processing or mixing steps performed to generate a desired output.

[0155] At block 1108, a playback device plays the generated media content. In various examples, this playback can be performed simultaneously with and / or in synchronization with additional playback devices.

[0156] FIG. 12 shows another example method 1200 for generating and playing generated media content. As described above, it can be beneficial to incorporate or rely on blockchain data to produce generated media content. Method 1200 begins, in block 1202, by accessing, via a playback device, blockchain data stored via a distributed ledger. The distributed ledger can be a public blockchain such as Ethereum, Bitcoin, Solana, etc., or optionally a private or semi-private blockchain, or a non-blockchain ledger. The blockchain data can include one or more existing media segments or other seeds used in the generation of the generated media content. In some examples, such data is stored directly on the blockchain itself, while in other examples, the blockchain can store a pointer (e.g., a URL or URI) indicating the location where the media segment or other seed data is stored. Optionally, the blockchain data takes the form of one or more non-fungible tokens (NFTs).

[0157] Method 1200, at block 1204, generates media content at least partially based on data of a blockchain via a playback device. In some examples, this generation includes accessing a library of existing media segments stored on the playback device or other suitable storage locations (e.g., a remote server, other devices on a local network, etc.). In some cases, these media segments are fetched from the blockchain or other remote locations and stored via the playback device. Next, the playback device can arrange a selection of existing media segments from the library for playback according to a generated media content model. This selection can be at least partially based on data of the blockchain. For example, as described elsewhere in this specification, the generated media content model can communicate with a smart contract or a decentralized autonomous organization (DAO) in a way that affects the specific generated media content output by the model. In some cases, specific NFTs or other tokens may affect the output of the generated media content model. Such NFTs or other tokens can, in some cases, be used in combination such that a specific combination of NFTs or other tokens generates a unique output via the generated media content model. In at least some instances, two or more blockchains can be utilized simultaneously in this way (e.g., the generated media content model can vary the output based on a user who holds both a first NFT on the Solana network and a second NFT on the Ethereum network). At block 1206, method 1200 includes playing back the generated media content via the playback device.

[0158] Figure 13 shows another exemplary method 1300 for generating and playing generated media content. As described above, smart contracts and other self-executing code can be used for generating, storing, and playing generated media content. In the first block 1302, method 1300 transmits, via a playback device and through a network, data associated with a first token to a network address of a distributed ledger. The address can be associated with a generated media smart contract configured to generate a generated media content model. In some examples, the first token can be an NFT, and optionally, multiple such tokens can be sent to the address of the smart contract.

[0159] In block 1304, method 1300 receives, via a playback device, a generated media content model from a network address associated with the generated media smart contract. For example, when the smart contract is executed, it generates a specific generated media content model based at least in part on the data associated with the first token. This generated media content model can be provided to the user, and as a result, new and original media content can be generated based on the first token data. In addition to the first token data, the smart contract can also generate different outputs based on other input parameters (such as sensor data, playback device characteristic data, playback device state, user listening history data, etc.) as described elsewhere, provided that such data is provided to the smart contract address. Additionally, or alternatively, the generated media content model provided to the user can output different media content based on one or more other input parameters.

[0160] Next, in block 1306, the playback device generates media content based on at least a portion of the generated media content model. In some examples, this generation includes accessing a library of existing media segments stored on the playback device or other suitable storage locations (e.g., a remote server, other devices on a local network, etc.). Next, the playback device can arrange the existing media segments selected from the library for playback according to the generated media content model. In block 1308, method 1300 plays the generated media content via the playback device.

[0161] Various examples of generated media playback are described herein. As will be appreciated by those skilled in the art, a wide variety of generated media modules, algorithms, inputs, sensor data, and playback device configurations are contemplated and can be used in accordance with this technology.

[0162] IV. Conclusion The foregoing discussion of playback devices, controller devices, playback zone configurations, and media content sources provides only some examples of operating environments in which the functions and methods described below can be implemented. Other operating environments and configurations of media playback systems, playback devices, and network devices not explicitly described herein are also applicable and may be suitable for implementing the functions and methods.

[0163] The foregoing description discloses, among other things, various exemplary systems, methods, apparatuses, and products that include firmware and / or software executed on hardware. It is understood that such examples are merely exemplary and should not be considered limiting. For example, any or all of the firmware, hardware, and / or software aspects or components can be implemented exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Thus, the examples provided are not the only way to implement such systems, methods, apparatuses, and / or products.

[0164] Furthermore, references to "examples" in this specification mean that a particular feature, structure, or characteristic described in connection with the example can be included in at least one example or embodiment of the invention. The appearance of this phrase in various places in this specification does not necessarily refer to the same example, nor is it an alternative or alternative example mutually exclusive with other examples. Thus, the examples described in this specification, as explicitly and implicitly understood by those skilled in the art, can be combined with other examples.

[0165] This specification is presented primarily with respect to other symbolic representations that are directly or indirectly similar in operation to an exemplary environment, system, procedure, step, logical block, process, and data processing device coupled to a network. The description and representation of these processes are typically used by those skilled in the art to most effectively convey the substance of their work to other skilled artisans. Numerous specific details are set forth in order to provide a complete understanding of the present disclosure. However, it will be understood by those skilled in the art that specific examples of the technology can be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail in order to avoid unnecessarily obscuring aspects of the examples. Thus, the scope of the present disclosure is defined by the appended claims rather than the foregoing description of the examples.

[0166] If any of the appended claims is read to cover purely software and / or firmware embodiments, at least one of the elements in at least one example is explicitly defined herein to include a tangible non-transitory medium such as a memory, DVD, CD, Blu-ray™, etc. that stores the software and / or firmware.

[0167] The disclosed technology is illustrated, for example, according to the various examples described below. The various examples of embodiments of the disclosed technology are described as numbered examples (1, 2, 3, etc.) for convenience. These are provided as examples and do not limit the disclosed technology. Note that any of the dependent examples may be combined in any combination or incorporated into each independent example. Other examples can be presented similarly.

[0168] Example 1: A method comprising receiving input parameters at a coordinator device, transmitting the input parameters from the coordinator device to a plurality of playback devices each having a generation media module internally, and transmitting timing data from the coordinator device to the plurality of playback devices such that the playback devices simultaneously play generation media content based at least in part on the input parameters.

[0169] Example 2: The method of any one of the examples herein, wherein a first and a second playback device each play different generated audio content based at least in part on input parameters.

[0170] Example 3: The input parameter is one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled to another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity), user mood data) in any one of the methods of the examples herein.

[0171] Example 4: The timing data includes at least one of clock data or one or more synchronization signals in any one of the methods of the examples herein.

[0172] Example 5: The method of any one of the examples herein further includes the step of transmitting, from a coordinator device, a signal for changing a generation media module of a playback device to at least one of a plurality of playback devices.

[0173] Example 6: The generated media content includes at least one of generated audio content or generated visual content in any one of the methods of the examples herein.

[0174] Example 7: The generation media module includes an algorithm that automatically generates a new media output based on an input including at least input parameters in any one of the methods of the examples herein.

[0175] Example 8: A device comprising a network interface, one or more processors, and a tangible non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the device to perform operations, the operations including receiving input parameters via the network interface, transmitting the input parameters via the network interface to a plurality of playback devices each having a generation media module therein, and transmitting timing data via the network interface to the plurality of playback devices such that the playback devices simultaneously play generation media content based at least in part on the input parameters.

[0176] Example 9: A device according to any one of the examples herein, wherein the first and second playback devices each play different generated audio content based at least in part on the input parameters.

[0177] Example 10: A device according to any one of the examples herein, wherein the input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device state (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled to another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity), user mood data).

[0178] Example 11: A device according to any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0179] Example 12: The operation further includes a step of transmitting, from a coordinator device to at least one of a plurality of playback devices, a signal that causes a generation media module of the playback device to be changed via a network interface, for any one device of the examples herein.

[0180] Example 13: The generated media content includes at least one of generated audio content or generated visual content, for any one device of the examples herein.

[0181] Example 14: The generation media module includes an algorithm that automatically generates a new media output based on an input that includes at least input parameters, for any one device of the examples herein.

[0182] A tangible non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a device, cause the device to perform an operation, the operation including receiving input parameters at a coordinator device; transmitting the input parameters from the coordinator device to a plurality of playback devices each having a generation media module therein; and transmitting timing data from the coordinator device to the plurality of playback devices such that the playback devices simultaneously play generated media content based at least in part on the input parameters.

[0183] Example 16: A first playback device and a second playback device each play different generated audio content based at least in part on input parameters, for any one computer-readable medium of the examples herein.

[0184] Example 17: A computer-readable medium of any one of the examples herein, wherein the input parameter includes one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled to another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity), user mood data).

[0185] Example 18: A computer-readable medium of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0186] Example 19: A computer-readable medium of any one of the examples herein, further comprising the step of transmitting, from a coordinator device, a signal for changing a generation media module of a playback device to at least one of a plurality of playback devices.

[0187] Example 20: A computer-readable medium of any one of the examples herein, wherein the generated media content includes at least one of generated audio content or generated visual content.

[0188] Example 21: A computer-readable medium of any one of the examples herein, wherein the generation media module includes an algorithm that automatically generates a new media output based on an input that includes at least input parameters.

[0189] Example 22: A method comprising: receiving input parameters in a coordinator device; generating first and second media content streams via a generation media module of the coordinator device; transmitting the first media content stream to a first playback device via the coordinator device; and transmitting the second media content stream to a second playback device via the coordinator device such that the first and second media content streams are simultaneously played back via the first and second playback devices.

[0190] Example 23: The method according to any one of the examples herein, further comprising transmitting timing data from the coordinator device to each of the first and second playback devices.

[0191] Example 24: The method according to any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0192] Example 25: The method according to any one of the examples herein, wherein the first and second media content streams are different.

[0193] Example 26: The input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled to another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity), user mood data). The method according to any one of the examples herein.

[0194] Example 27: A method according to any one of the examples herein, further comprising the step of modifying a generation media module of a coordinator device.

[0195] Example 28: A method according to any one of the examples herein, wherein each of the first and second generated media content streams includes at least one of generated audio content or generated visual content.

[0196] Example 29: A method according to any one of the examples herein, wherein the generation media module includes an algorithm that automatically generates a new media output based on an input that includes at least input parameters.

[0197] Example 30: A device comprising a network interface, a generation media module, one or more processors, and a tangible non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the device to perform operations including receiving input parameters via the network interface, generating first and second media content streams via the generation media module, transmitting the first media content stream to a first playback device via the network interface, and transmitting the second media content stream to a second playback device via the network interface such that the first and second media content streams are simultaneously played back via the first and second playback devices.

[0198] Example 31: A device according to any one of the examples herein, wherein the operations further include transmitting timing data to each of the first and second playback devices via the network interface.

[0199] Example 32: A device according to any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0200] Example 33: A device according to any one of the examples herein, wherein the first and second media content streams are different.

[0201] Example 34: An input parameter is one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device state (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled to another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity), user mood data). A device according to any one of the examples herein.

[0202] Example 35: An operation further includes the step of changing a generation media module. A device according to any one of the examples herein.

[0203] Example 36: Each of the first and second generated media content streams includes at least one of generated audio content or generated visual content. A device according to any one of the examples herein.

[0204] Example 37: A generation media module includes an algorithm that automatically generates a new media output based on an input including at least an input parameter. A device according to any one of the examples herein.

[0205] Example 38: A tangible non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a coordinator device, cause the coordinator device to perform operations, the operations including receiving input parameters at the coordinator device; generating first and second media content streams via a generation media module of the coordinator device; transmitting the first media content stream to a first playback device via the coordinator device; and transmitting the second media content stream to a second playback device via the coordinator device such that the first and second media content streams are played back simultaneously via the first and second playback devices.

[0206] Example 39: The computer-readable medium of any one of the examples herein, further including transmitting timing data from the coordinator device to each of the first and second playback devices.

[0207] Example 40: The computer-readable medium of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0208] Example 41: The computer-readable medium of any one of the examples herein, wherein the first and second media content streams are different.

[0209] Example 42: The input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled to another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity), user mood data), in any one of the examples of this specification, a computer-readable medium.

[0210] Example 43: The operation further includes the step of changing the generation media module of the coordinator device, in any one of the examples of this specification, a computer-readable medium.

[0211] Example 44: Each of the first and second generated media content streams includes at least one of generated audio content or generated visual content, in any one of the examples of this specification, a computer-readable medium.

[0212] Example 45: The generation media module includes an algorithm that automatically generates a new media output based on an input that includes at least input parameters, in any one of the examples of this specification, a computer-readable medium.

[0213] Example 46: A playback device, comprising: one or more amplifiers configured to drive one or more audio transducers; one or more processors; and a data storage device having instructions that, when executed by the one or more processors, cause the playback device to perform operations, the operations including: receiving, at the playback device, one or more first input parameters; generating, at least in part based on the one or more first input parameters via the playback device, first media content, the generating including accessing a library stored in the playback device that includes a plurality of existing media segments, and arranging a first selection of existing media segments from the library for playback based at least in part on the one or more input parameters according to a generated media content model; and playing back the first generated media content via the one or more amplifiers.

[0214] Example 47: The operations include: receiving, at the playback device, one or more second input parameters different from the one or more first input parameters; generating, at least in part based on the one or more second input parameters via the playback device, second media content, the second media content being different from the first media content, the generating including accessing the library and arranging a second selection of existing media segments from the library for playback based at least in part on the one or more second input parameters according to a generated media content model; and playing back the second generated media content via the one or more amplifiers, wherein the playback device is any one of the examples herein.

[0215] Example 48: The playback device of claim 1, wherein arranging a first selection of existing media segments from the library for playback includes arranging at least partially temporally offset two or more of the existing media segments.

[0216] Example 49: A playback device according to any one of the examples herein, wherein the step of arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments to at least partially temporally overlap.

[0217] Example 50: A playback device according to any one of the examples herein, wherein the step of arranging a first selection of existing media segments from a library for playback includes applying different equalization adjustments to different existing media segments.

[0218] Example 51: A playback device according to any one of the examples herein, wherein the step of arranging a first selection of existing media segments from a library or for playback includes applying gain levels that vary over time to different existing media segments.

[0219] Example 52: A playback device according to any one of the examples herein, wherein the step of arranging a first selection of existing media segments from a library or for playback includes randomizing the start point for the playback of a particular existing media segment.

[0220] Example 53: A playback device according to any one of the examples herein, wherein the first generated media content and the second generated media content each include new media content.

[0221] Example 54: A playback device according to any one of the examples herein, wherein the first generated media content includes audio content and the plurality of existing media segments include a plurality of existing audio segments.

[0222] Example 55: A playback device according to any one of the examples herein, wherein the first generated media content includes audio-visual content and the plurality of existing media segments include a plurality of existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0223] Example 56: A playback device according to any one of the examples herein, further comprising receiving, via a network interface, additional existing media segments, and updating a library to include at least the additional existing media segments.

[0224] Example 57: The first and second input parameters are physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device state (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled to another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity, speech characteristics), user mood data), and a playback device according to any one of the examples herein includes one or more of the foregoing.

[0225] Example 58: A method comprising receiving, at a playback device, one or more first input parameters; generating, via the playback device, first media content based at least in part on the one or more first input parameters, the generating step including accessing a library stored in the playback device that includes a plurality of existing media segments, and arranging a first selection of existing media segments from the library for playback based at least in part on the one or more input parameters according to a generated media content model; and playing back the first generated media content via the playback device.

[0226] Example 59: A method according to any one of the examples herein, comprising: receiving, at a playback device, one or more second input parameters different from a first input parameter; generating, via the playback device, second media content based at least in part on the one or more second input parameters, the second media content being different from first media content, the generating step including accessing a library and arranging a second selection of existing media segments from the library for playback based at least in part on the one or more second input parameters according to a generated media content model; and playing back, via the playback device, the second generated media content.

[0227] Example 60: A method according to any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging at least a portion of two or more of the existing media segments to be at least partially temporally offset.

[0228] Example 61: A method according to any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging at least a portion of two or more of the existing media segments to be at least partially temporally overlapping.

[0229] Example 62: A method according to any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes applying different equalization adjustments to different existing media segments.

[0230] Example 63: A method according to any one of the examples herein, wherein arranging a first selection of existing media segments from a library or playback includes applying gain levels that vary over time to different existing media segments.

[0231] Example 64: The method of any one of the examples herein, wherein the step of arranging a first selection of an existing media segment from a library or playback includes randomizing a starting point for playback of a particular existing media segment.

[0232] Example 65: The method of any one of the examples herein, wherein the first generated media content and the second generated media content each include new media content.

[0233] Example 66: The method of any one of the examples herein, wherein the first generated media content includes audio content and the plurality of existing media segments include a plurality of existing audio segments.

[0234] Example 67: The method of any one of the examples herein, wherein the first generated media content includes audio-visual content and the plurality of existing media segments include a plurality of existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0235] Example 68: The method of any one of the examples herein, further comprising receiving additional existing media segments via a network interface and updating a library to include at least the additional existing media segments.

[0236] Example 69: The first and second input parameters are physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device state (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled to another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity, speech characteristics), user mood data), one or more of which are included in any one of the methods of the examples herein.

[0237] Example 70: A tangible non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a playback device, cause the playback device to perform operations, the operations including receiving, at the playback device, one or more first input parameters; generating, via the playback device, first media content based at least in part on the one or more first input parameters, the generating including accessing a library stored in the playback device that includes a plurality of existing media segments, and arranging a first selection of existing media segments from the library for playback based at least in part on the one or more input parameters according to a generated media content model; and playing back, via the playback device, the first generated media content.

[0238] Example 71: The operation includes receiving, at a playback device, one or more second input parameters that are different from a first input parameter, and generating second media content at least partially based on the one or more second input parameters via the playback device, wherein the second media content is different from first media content, and the generating step includes accessing a library and arranging a second selection of existing media segments from the library for playback at least partially based on the one or more second input parameters according to a generated media content model, and playing the second generated media content via one or more amplifiers, a computer-readable medium of any one of the examples herein.

[0239] Example 72: The step of arranging a first selection of existing media segments from a library for playback includes arranging at least two of the existing media segments to be at least partially temporally offset, a computer-readable medium of any one of the examples herein.

[0240] Example 73: The step of arranging a first selection of existing media segments from a library for playback includes arranging at least two of the existing media segments to be at least partially temporally overlapping, a computer-readable medium of any one of the examples herein.

[0241] Example 74: The step of arranging a first selection of existing media segments from a library for playback includes applying different equalization adjustments to different existing media segments, a computer-readable medium of any one of the examples herein.

[0242] Example 75: The step of arranging a first selection of existing media segments from a library for playback includes applying gain levels that vary over time to different existing media segments, a computer-readable medium of any one of the examples herein.

[0243] Example 76: A computer-readable medium according to any one of the examples herein, wherein the step of placing a first selection of an existing media segment from a library for reproduction includes the step of randomizing a start point for reproduction of a particular existing media segment.

[0244] Example 77: A computer-readable medium according to any one of the examples herein, wherein the first generated media content and the second generated media content each include new media content.

[0245] Example 78: A computer-readable medium according to any one of the examples herein, wherein the first generated media content includes audio content and the plurality of existing media segments include a plurality of existing audio segments.

[0246] Example 79: A computer-readable medium according to any one of the examples herein, wherein the first generated media content includes audio-visual content and the plurality of existing media segments include a plurality of existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0247] Example 80: A computer-readable medium according to any one of the examples herein, further comprising the step of receiving additional existing media segments via a network interface and the step of updating the library to include at least the additional existing media segments.

[0248] Example 81: The first and second input parameters are physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled to another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity, speech characteristics), user mood data), a computer-readable medium of any one of the examples herein containing one or more of these.

[0249] Example 82: A system comprising a first playback device and a second playback device. The first playback device comprises a first network interface, one or more first processors, and a data storage device having instructions that, when executed by the one or more processors, cause the first playback device to perform operations, the operations including receiving one or more input parameters, generating media content based at least in part on the one or more input parameters, the generated media content including a first portion and at least a second portion, the generating step including accessing a library stored in the playback device that includes a plurality of existing media segments, and arranging for the selection of existing media segments from the library for playback based at least in part on the one or more input parameters according to a media content model, transmitting a signal including the second portion of the generated media content and corresponding timing information via the first network interface, and causing the playback of the first portion of the generated media content. The second playback device comprises a second network interface, one or more audio transducers, one or more second processors, and a data storage device having instructions that, when executed by the one or more second processors, cause the second playback device to perform operations, the operations including receiving a signal transmitted from the first playback device via the second network interface, and playing back the second portion of the generated media content according to the timing information substantially in synchronization with the playback of the first portion of the generated media content via the one or more transducers.

[0250] Example 83: The system further includes a network device, which includes a third network interface, one or more processors, and a data storage device having instructions that, when executed by the one or more processors, cause the third playback device to perform operations. The operations include receiving a request from the first playback device via the third network interface across a data network, and in response to receiving the request, transmitting an updated library of existing media segments to the first playback device via the third network interface across the data network. The system is any one of the examples described herein.

[0251] Example 84: The network device includes one or more of a remote server, another playback device, a mobile computing device, a laptop, or a tablet. The system is any one of the examples described herein.

[0252] Example 85: The system includes a first playback device and a second playback device communicatively coupled via a local area network. The first playback device includes one or more first processors, one or more first audio transducers, and a data storage device having instructions that, when executed by the one or more first processors, cause the first playback device to perform operations. The operations include receiving one or more input parameters, generating first media content based at least in part on the one or more input parameters, including accessing a first library stored in the first playback device that includes a plurality of existing media segments, and arranging a selection of existing media segments from the first library for playback based at least in part on the one or more input parameters according to a first generated media content model, and playing the first generated media content via the one or more first audio transducers. The second playback device includes a second network interface.

[0253] One or more second audio transducers, one or more second processors, and a data storage device having instructions that, when executed by the one or more second processors, cause an operation to be performed on a second playback device, the operation comprising: generating second media content based at least in part on one or more input parameters, wherein the second generated media content is substantially identical to the first generated media content, the generating step comprising accessing a second library stored on the second playback device that includes a plurality of existing media segments, and arranging a selection of existing media segments from the second library for playback based at least in part on the one or more input parameters according to a second generated media content model; and playing the second generated media content in synchronization with the playing of the first generated media content via the first playback device via the one or more second audio transducers. A system.

[0254] Example 86: A system according to any one of the examples herein, wherein the first generated media content model and the second generated media content model are substantially identical.

[0255] Example 87: A system according to any one of the examples herein, wherein the first library and the second library are substantially identical.

[0256] Example 88: A method comprising the following steps: accessing, via a playback device, data of a blockchain stored on a distributed ledger; generating, via the playback device, media content based at least in part on the data of the blockchain; wherein the generating step comprises: accessing a library stored on the playback device and containing a plurality of existing media segments, and arranging a selection of existing media segments from the library for playback according to a generated media content model and based at least in part on the data of the blockchain; and further, playing back the generated media content via the playback device.

[0257] Example 89: A method according to any one of the examples herein, wherein the NFT data includes one or more existing media segments, and accessing the NFT data includes storing the one or more existing media segments in a library.

[0258] Example 90: A method according to any one of the examples herein, wherein the data of the blockchain includes first non-fungible token (NFT) data, wherein the distributed ledger is a first distributed ledger, and the method further includes accessing, via a playback device, data related to a second NFT stored on a second distributed ledger, wherein arranging a selection of existing media segments from the library for playback according to a generated media content model is based at least in part on both the first NFT data and the second NFT data.

[0259] Example 91: A method according to any one of the examples herein, wherein the first distributed ledger is associated with a first blockchain layer, and the second distributed ledger is associated with a second blockchain layer different from the first blockchain layer.

[0260] Example 92: A method according to any one of the examples herein, wherein the data of the blockchain is associated with a playlist.

[0261] Example 93: A method according to any one of the examples herein, wherein blockchain data depends at least in part on a transaction recorded in a distributed ledger that includes non-fungible tokens (NFTs).

[0262] Example 94: Arranging the selection of existing media segments from a library according to a generative media content model further depends at least in part on one or more input parameters, a method according to any one of the examples herein.

[0263] Example 95: A method according to any one of the examples herein, wherein the input parameter includes one or more of physiological sensor data, networked device sensor data, environmental data, playback device characteristic data, playback device state, user listening history data, oracle data stored via a distributed ledger, or user data.

[0264] Example 96: A method according to any one of the examples herein, wherein user listening history data is stored via a distributed ledger.

[0265] Example 97: A method according to any one of the examples herein, wherein accessing blockchain data includes connecting to a user wallet that holds non-fungible tokens (NFTs).

[0266] Example 98: A method according to any one of the examples herein, wherein accessing blockchain data includes accessing code associated with a physical media object via a control device (e.g., a QR code or other code printed on custom vinyl or other media).

[0267] Example 99: A method according to any one of the examples herein, wherein arranging the selection of existing media segments from a library for playback includes arranging two or more of the existing media segments in a temporally offset manner.

[0268] Example 100: The method of any one of the examples herein, wherein arranging the selection of existing media segments from a library for regeneration includes arranging two or more of the existing media segments to at least partially overlap in time.

[0269] Example 101: A method comprising the following steps: transmitting, via a playback device and via a network, data associated with a first token to a network address of a distributed ledger, the address being associated with a generative media smart contract configured to generate a generative media content model; receiving, via the playback device, the generative media content model from the network address associated with the generative media smart contract; generating media content via the playback device at least partially based on the generative media content model; wherein the generating step includes accessing a library containing a plurality of existing media segments; and arranging a selection of existing media segments from the library for playback according to the generative media content model; the method further comprising playing back the generated media content via the playback device.

[0270] Example 102: The token data includes first non-fungible token (NFT) data, and the method further includes transmitting, via a playback device, data related to a second NFT stored in a distributed ledger to a network address associated with a generated media smart contract; receiving, via the playback device, a second generated media content model from a network address associated with the generated media smart contract, where the second generated media content model is different from the first generated media content model; generating, via the playback device, second media content at least partially based on the second generated media content model; and playing, via the playback device, the second generated media content; the method according to any one of the examples herein.

[0271] Example 103: The method according to any one of the examples herein, wherein the token data is associated with a curated playlist.

[0272] Example 104: The method according to any one of the examples herein, wherein the token data depends at least in part on a transaction recorded in a distributed ledger containing tokens.

[0273] Example 105: The step of arranging the selection of existing media segments from a library according to a generated media content model is further based at least in part on one or more input parameters; the method according to any one of the examples herein.

[0274] Example 106: The method according to any one of the examples herein, wherein the input parameters include one or more of physiological sensor data, networked device sensor data, environmental data, playback device characteristic data, playback device state, user listening history data, oracle data stored via a distributed ledger, or user data.

[0275] Example 107: The method according to any one of the examples herein, wherein the user's listening history data is stored via a distributed ledger.

[0276] Example 108: Any one of the methods of the present specification, further comprising accessing token data by connecting to a user wallet storing the token data before transmitting the token data.

[0277] Example 109: Any one of the methods of the present specification, further comprising accessing token data via a code associated with a physical media object (e.g., a QR code or other code printed on custom vinyl or other media) via a control device before transmitting the token data.

[0278] Example 110: Any one of the methods of the present specification, wherein the step of arranging the selection of existing media segments from a library for playback includes arranging two or more of the existing media segments in a manner that is at least partially temporally offset.

[0279] Example 111: Any one of the methods of the present specification, wherein the step of arranging the selection of existing media segments from a library for playback includes arranging two or more of the existing media segments in a manner that is at least partially temporally overlapping.

[0280] Example 112: Any one of the methods of the present specification, wherein the first generated media content and the second generated media content each include new media content.

[0281] Example 113: One or more tangible, non-transitory computer-readable media storing instructions executable by one or more processors to cause a media playback system or playback device to perform operations according to any one of the methods of the present specification.

[0282] Example 114: A media playback system comprising one or more processors; and a computer-readable medium holding any one of the methods of the present specification.

[0283] Example 115: A playback device comprising one or more processors; and a computer-readable medium holding any one of the methods of the examples herein.

Claims

1. In a coordinator device, receiving input parameters; transmitting the input parameters from the coordinator device to a plurality of playback devices each having a generation media module therein; transmitting timing data from the coordinator device to the plurality of playback devices so that the playback devices simultaneously play generation media content based at least in part on the input parameters; A method comprising:

2. The method according to claim 1, wherein a first and a second playback device each play different generated audio content based at least in part on the input parameters.

3. In a coordinator device, receiving input parameters; generating first and second media content streams via a generation media module of the coordinator device; transmitting the first media content stream to a first playback device via the coordinator device; transmitting the second media content stream to a second playback device via the coordinator device so that the first and second media content streams are simultaneously played via the first and second playback devices; A method comprising:

4. The method according to claim 3, further comprising transmitting timing data from the coordinator device to each of the first and second playback devices.

5. The method according to claim 4, wherein the first and second media content streams are different.

6. The method according to claim 4, wherein the timing data includes at least clock data or one or more synchronization signals.

7. The method according to any one of claims 3 to 6, wherein each of the first and second generated media content streams includes at least one of generated audio content or generated visual content.

8. The method according to any one of claims 3 to 7, wherein the generation media module includes an algorithm that automatically generates a new media output based on an input including at least the input parameters.

9. Receiving one or more first input parameters in a playback device; Generating first media content at least partially based on the one or more first input parameters via the playback device; Comprising, wherein the generating step Accessing a library stored in the playback device that includes a plurality of existing media segments; Arranging a first selection of existing media segments from the library for playback based at least partially on the one or more input parameters according to a generated media content model; Playing the generated first media content via the playback device; A method comprising.

10. The method according to claim 9, further comprising Receiving, in the playback device, one or more second input parameters different from the first input parameters; Generating second media content at least partially based on the one or more second input parameters via the playback device, wherein the second media content is different from the first media content; comprising, wherein the generating step comprises: accessing the library; arranging a second selection of existing media segments from the library for playback based at least in part on the one or more second input parameters according to the generated media content model; playing the generated second media content via the playback device; and including. **Claim 11** The method according to claim 10, wherein the generated first media content and the generated second media content each include new media content. **Claim 12** The method according to any one of claims 9 to 11, wherein the generated first media content includes audio content and the plurality of existing media segments include a plurality of existing audio segments. **Claim 13** The method according to any one of claims 9 to 12, wherein the generated first media content includes audio-visual content and the plurality of existing media segments include a plurality of existing audio segments, existing visual media segments, or existing audio-visual media segments. **Claim 14** receiving further existing media segments via a network interface; updating the library to include at least the further existing media segments; The method according to any one of claims 9 to 13, further comprising. **Claim 15** A playback device, one or more amplifiers configured to drive one or more audio transducers; one or more processors; A data storage device having instructions that, when executed by the one or more processors, cause the playback device to execute the method according to any one of claims 9 to 14; A playback device comprising the same. **Claim 16** A first playback device based on the playback device according to claim 15; A second network interface, one or more audio transducers, and a second playback device having one or more second processors; A system comprising: The second processor: Receives, via the second network interface, a signal transmitted from the first playback device having a second portion of the generated media content and timing information; Plays back the second portion of the generated media content according to the timing information via the one or more transducers so as to be substantially synchronized with the playback of the first portion of the generated media content via the first playback device; A system. **Claim 17** The system according to claim 16, further comprising a network device, the network device having: A third network interface; One or more processors; and A data storage device storing instructions executable by the one or more processors, wherein when the instructions are executed by the one or more processors, the third playback device: Receives a request from the first playback device via the third network interface of the data network; In response to receiving the request, transmits an updated library for the existing media segment to the first playback device via the third network interface of the data network. **Claim 18** The system according to claim 17, wherein the network device has one or more remote servers, other playback devices, mobile computing devices, laptops, or tablets.

19. A first playback device based on the playback device according to claim 15, and A second playback device having a second network interface, one or more second audio transducers, one or more second processors, and a data storage device storing instructions executable by the one or more second processors, and A system comprising: Wherein when the instructions are executed by the one or more second processors, the second playback device: By accessing a second library stored in the second playback device containing a plurality of existing media segments, and In accordance with a second generated media content model, by arranging a selection of existing media segments from the second library for playback based at least in part on one or more input parameters, Generating second media content based at least in part on one or more input parameters: and Via the one or more second audio transducers, playing the generated second media content in synchronization with the playing of the generated first media content via the first playback device. System.

20. The system according to claim 19, wherein the first generated media content model and the second generated media content model are substantially identical.

21. The system according to claim 19 or 20, wherein the first library and the second library are substantially identical.

22. The method according to claims 1 to 14, wherein receiving one or more input parameters Accessing data of a blockchain stored in a distributed ledger via a playback device, where the first one or more parameters are composed of the data of the blockchain. ーn data, where the first one or more parameters are composed of the data of the blockchain.

23. The method according to claim 22, wherein the data of the blockchain is composed of one or more existing media segments, and accessing the data of the blockchain includes storing one or more existing media segments in a library.

24. The method according to claim 22 or 23, wherein the data of the blockchain includes first non-fungible token (NFT) data, where the distributed ledger is a first distributed ledger, and further, the method includes accessing data related to a second NFT stored in a second distributed ledger via a playback device, and also, where arranging the selection of existing media segments from the library for playback according to a generated media content model is at least partially based on both the first NFT data and the second NFT data.

25. The method according to any one of claims 22 to 24, wherein the first distributed ledger is associated with a first blockchain layer, and also, where the second distributed ledger is associated with a second blockchain layer different from the first blockchain layer.

26. The method according to any one of claims 22 to 25, wherein the data of the blockchain is associated with a playlist.

27. The method according to any one of claims 22 to 25, wherein at least a part of the data of the blockchain depends on a transaction recorded in a distributed ledger including non-fungible tokens (NFTs).

28. The method according to any one of claims 22 to 25, wherein arranging the selection of existing media segments from a library according to a generated media content model is further based at least in part on one or more additional input parameters.

29. The method according to any one of claims 22 to 28, wherein access to the data of the blockchain is made by connecting to a user wallet holding non-fungible tokens (NFTs).

30. The method according to any one of claims 22 to 29, wherein access to the data of the blockchain is made by accessing the code associated with a physical media object via a control device.

31. A step of transmitting, by a playback device, data associated with a first token to a network address of a distributed ledger through a network, wherein this address is associated with a generated media smart contract configured to generate a generated media content model; A step of receiving, by a playback device, a generated media content model from a network address associated with a generated media smart contract; A step of generating media content by a playback device based at least in part on the generated media content model, comprising The step of generating comprises Accessing a library consisting of a plurality of existing media segments; Arranging a selection of existing media segments from the library for playback according to a generated media content model; having Furthermore, a step of playing back the generated media content by a playback device, Method.

32. The method according to claim 31, wherein the token data is composed of first non-fungible token (NFT) data, Further, the method comprises: transmitting, via a playback device, data related to a second NFT stored in a distributed ledger to a network address related to a generation media smart contract; receiving, by the playback device, a second generation media content model from a network address associated with the generation media smart contract, wherein the second generation media content model is different from the first generation media content model; generating, by the playback device, second media content based at least in part on the second generation media content model; playing, by the playback device, the second generated media content; and comprising.

33. The method according to claim 31 or 32, wherein the token data is associated with a curated playlist.

34. The method according to any one of claims 31 to 33, wherein the token data depends at least in part on a transaction recorded in a distributed ledger in which the token is involved.

35. The method according to any one of claims 31 to 34, wherein arranging the selection of existing media segments from a library according to a generation media content model is further based at least in part on one or more input parameters.

36. The method according to any one of claims 31 to 35, further comprising connecting to a user wallet storing the token data to access the token data before transmitting the token data.

37. The method according to any one of claims 31 to 36, further comprising the step of accessing the token data via a code associated with a physical media object through a control device before transmitting the token data.

38. The method according to any one of claims 31 to 36, wherein the first and / or second input parameters are composed of one or more of the following: Physiological sensor data Sensor data of network devices Environmental data Playback device capability data The state of the playback device User data Oracle data stored via a distributed ledger.

39. The method according to any one of claims 22 to 38, wherein arranging the first selection of existing media segments from the library for playback includes at least one of the following: Arranging two or more existing media segments at least partially temporally offset; Arranging two or more existing media segments at least partially temporally overlapping; Applying different equalization adjustments to different existing media segments; Applying gain levels that change over time to different existing media segments; Randomizing the playback start point of a specific existing media segment.

40. The method according to any one of claims 22 to 39, wherein the user's viewing history data is stored via a distributed ledger.

41. The method according to any one of claims 22 to 40, wherein the first generated media content and the second generated media content are each composed of new media content.

42. One or more tangible, non-transitory computer-readable media storing instructions executable by one or more processors to cause a media playback system or a playback device to perform the operation according to any one of claims 1 to 14 or 22 to 41. **Claim 43** A media playback system One or more processors and The computer-readable medium according to claim 42. **Claim 44** A playback device One or more processors and The computer-readable medium according to claim 42.

Citation Information

Patent Citations

  • Music content generation

    WO2021163377A1