Digital media generation based on blockchain data

The media playback system generates personalized and adaptive audio experiences by using algorithms and contextual data to create dynamic media content, addressing the limitations of traditional systems and incorporating blockchain for secure content creation.

JP7797707B2Active Publication Date: 2026-01-13SONOS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024568610
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-18
Filing Date
2023-05-09
Publication Date
2026-01-13
Estimated Expiration
2043-05-09

AI Technical Summary

Technical Problem

Existing digital audio playback systems lack the ability to dynamically generate and modify media content in response to real-time user inputs and environmental factors, limiting the personalization and adaptability of the listening experience.

Method used

A media playback system that utilizes generative media content, which is synthesized based on algorithms and contextual data, including user inputs, sensor data, and environmental conditions, and can be stored on blockchain databases via non-fungible tokens (NFTs) to create dynamic and personalized audio experiences.

Benefits of technology

Enables unique and adaptive user experiences by dynamically generating media content that can change in real-time, enhancing user engagement and personalization based on emotional states, location, activity, and other factors, and integrating with blockchain technology for secure media content creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007797707000001
    Figure 0007797707000001
  • Figure 0007797707000002
    Figure 0007797707000002
  • Figure 0007797707000003
    Figure 0007797707000003
Patent Text Reader

Abstract

Generated media content (e.g., generated audio) can be dynamically generated based on various inputs that can include blockchain data. A playback device accesses blockchain data stored via a distributed ledger and generates media content based on at least a portion of the blockchain data. The playback device can access a library of existing media segments and, according to a generated media content model, place a selection of existing media segments from the library for playback based at least in part on the blockchain data. The generated media content can be played back via the playback device.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Application No. 63 / 364,931, filed May 18, 2022, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to consumer goods, and more particularly to methods, systems, products, features, services, and other elements directed to media playback or some aspect thereof. [Background technology]

[0003] Options for accessing and listening to digital audio at high volume settings were limited until 2002, when SONOS, Inc. began developing a new type of playback system. Sonos then filed one of the first patent applications, titled "Method for Synchronizing Audio Playback between Multiple Networked Devices," in 2003 and began offering its first commercially available media playback system in 2005. The Sonos wireless home sound system allows people to experience music from many sources through one or more networked playback devices. Through a software control application installed on a controller (e.g., smartphone, tablet, computer, audio input device), users can play what they want in any room with a networked playback device. Media content (e.g., songs, podcasts, video sounds) can be streamed to the playback devices so that each room with a playback device can play corresponding different media content. Furthermore, rooms can be grouped together for synchronized playback of the same media content and / or the same media content can be listened to synchronously in all rooms. [Brief explanation of the drawings]

[0004] The features, aspects, and advantages of the disclosed technology may be better understood with regard to the following description, appended claims, and accompanying drawings, set forth below. Those skilled in the art will recognize that the features shown in the drawings are for illustrative purposes and that variations, including different and / or additional features and arrangements thereof, are possible. [Figure 1A] 1 is a partial cutaway view of an environment having a media playback system configured in accordance with aspects of the disclosed technology. [Figure 1B] 1B is a schematic diagram of the media playback system and one or more networks of FIG. [Figure 1C] Playback device block diagram [Figure 1D] Playback device block diagram [Figure 1E] Block diagram of a combined playback device [Figure 1F] Network Microphone Device Block Diagram [Figure 1G] Playback device block diagram [Figure 1H] Partial schematic diagram of the control device [Figure 1I] Schematic diagram of supported media playback system zones [Figure 1J] Schematic diagram of supported media playback system zones [Figure 1K] Schematic diagram of supported media playback system zones [Figure 1L] Schematic diagram of supported media playback system zones [Figure 1M] Media Playback System Area Schematic [Figure 2] 1 is a functional block diagram of a system for playback of generated media content according to an example of the present technology; [Figure 3] FIG. 1 is a functional block diagram of a generated media module according to aspects of the present technology. [Figure 4] FIG. 1 illustrates an example architecture for storing and retrieving generated media content in accordance with aspects of the present technology. [Figure 5]FIG. 1 is a functional block diagram illustrating data exchange in a system for playback of generated media content in accordance with aspects of the present technology. [Figure 6] FIG. 1 is a schematic diagram of an exemplary distributed generative media playback system in accordance with aspects of the present technology. [Figure 7] FIG. 1 is a schematic diagram of another exemplary distributed generative media playback system in accordance with aspects of the present technology. [Figure 8] 1 is a flow diagram of a method for playing generated media content in accordance with aspects of the present technology; [Figure 9] 1 is a flow diagram of a method for playing generated media content in accordance with aspects of the present technology; [Figure 10] 1 is a flow diagram of a method for playing generated media content in accordance with aspects of the present technology; [Figure 11] 1 is a flow diagram of a method for playing generated media content in accordance with aspects of the present technology; [Figure 12] 1 is a flow diagram of a method for playing generated media content in accordance with aspects of the present technology; [Figure 13] 1 is a flow diagram of a method for playing generated media content in accordance with aspects of the present technology;

[0005] The drawings are for purposes of illustrating examples of the technology; however, one skilled in the art would understand that the technology disclosed herein is not limited to the arrangements and / or instrumentality shown in the drawings. DETAILED DESCRIPTION OF THE INVENTION

[0006] I. Overview Generative media content is content that is dynamically synthesized, created, and / or modified based on algorithms, whether implemented in software or physical models. Generative media content can change over time based solely on algorithms or in conjunction with contextual data (e.g., user sensor data, environmental sensor data, occurrence data). In various examples, such generative media content can include generative audio (e.g., music, ambient sounds, etc.), generative visual images (e.g., abstract visual designs that dynamically change lighting, shape, color, etc.), generative scents, generative haptic outputs (vibrations, haptic outputs, etc.), or any other suitable media content or combinations thereof. As described elsewhere herein, generative media can be generated, at least in part, via algorithms and / or non-human systems that utilize rule-based computations to generate new media content.

[0007] Because generated media content can change dynamically in real time, it enables unique user experiences not available using traditional media playback of pre-recorded content. For example, the generated audio can be endless and / or dynamic audio that changes as inputs to the algorithm (e.g., input parameters related to user input, sensor data, media source data, or any other suitable input data) change. In some examples, the generated audio can be used to steer a user's mood toward a desired emotional state, with one or more characteristics of the generated audio changing in response to real-time measurements that reflect the user's emotional state. As used in examples of the present technology, a system can provide generated audio based on the user's current and / or desired emotional state, based on the user's activity level, based on the number of users present in the environment, or based on any other suitable input parameters.

[0008] As another example, the generated audio can be created and / or modified based on one or more inputs, such as a user's location or activity, the number of users present in the room, the time of day, or any other input (e.g., as determined by one or more sensors or user input). For example, when one user is sitting at their desk in a calm state, a media playback system can automatically generate generated audio content suitable for intensive study or work, while when multiple users are present in the room in an excited state with a lot of movement, the same media playback system can automatically generate generated audio suitable for a social gathering or dance party. In various examples, audio characteristics that can be dynamically modified to produce the generated audio can include audio sample or clip selection, tempo, bass / treble / mid-range volume, spatial filtering of the audio output, or any other suitable audio characteristics. Audio characteristics can be modified by using audio samples that may have different tones or sounds, timing of tones or sounds, and / or desired qualities. In some cases, the playback of content can also be modified by filtering or modulating characteristics, such as equalization, phase, or reverb / delay. During the listening experience, the audio characteristics of the generated music can be altered based on several inputs, such as the time of day, geographic location, weather, or various user inputs, such as inferred mood, collective level of activity, or physiological inputs such as heart rate.

[0009] In some cases, generative media content, such as soundscapes, can be associated with data stored on one or more blockchain databases and / or layers. For example, non-fungible tokens (NFTs) typically consist of data files stored on a blockchain associated with a digital or tangible asset (e.g., a song or album, visual artwork, or literary work). For example, while an NFT may reference a digital artwork, it is not common for the NFT itself to contain the digital artwork itself because the required file size would be unwieldy (or too costly) to store on a blockchain layer. Instead, an NFT may consist of metadata (e.g., a URL or other locator) that indicates where the digital artwork is located and / or how to access the artwork. In various instances, NFTs and other data stored via blockchain or other distributed ledger technologies can be used in media content creation. For example, in some instances, NFTs serve as inputs to a generative media engine, resulting in generative media content with characteristics that depend, at least in part, on the particular NFT or other blockchain data.

[0010] Although some examples described herein may refer to functions performed by given parties, such as "users," "listeners," and / or other entities, it should be understood that this is for illustrative purposes only. The claims should not be construed as requiring action by any such example actors unless expressly required by the language of the claims themselves.

[0011] In the figures, like reference numbers generally indicate similar and / or identical elements. To facilitate the description of any particular element, the most significant digit(s) of the reference number refers to the figure in which that element is first introduced. For example, element 110a is first introduced and described with reference to FIG. 1A. Many of the details, dimensions, angles, and other features shown in the figures are merely illustrative of particular examples of the disclosed technology. Thus, other examples can have other details, dimensions, angles, and features without departing from the spirit or scope of the present disclosure. Moreover, those skilled in the art will recognize that further examples of the various disclosed technologies can be practiced without some of the details described below.

[0012] II. Appropriate Operating Environment 1A is a partial cutaway view of a media playback system 100 distributed within an environment 101 (e.g., a home). Media playback system 100 includes one or more playback devices 110 (individually identified as playback devices 110a-110n), one or more network microphone devices (“NMDs”) 120 (individually identified as NMDs 120a-120c), and one or more control devices 130 (individually identified as control devices 130a and 130b).

[0013] As used herein, the term "playback device" may generally refer to a network device configured to receive, process, and / or output data for a media playback system. For example, a playback device may be a network device that receives and processes audio content. In some examples, a playback device includes one or more transducers or speakers powered by one or more amplifiers. However, in other examples, a playback device includes one or neither of a speaker and an amplifier. For example, a playback device may include one or more amplifiers configured to drive one or more speakers external to the playback device via corresponding wires or cables.

[0014] Additionally, as used herein, the term NMD (i.e., "network microphone device") may generally refer to a network device configured for audio detection. In some examples, an NMD is a standalone device configured primarily for audio detection. In other examples, an NMD is incorporated into a playback device (or vice versa).

[0015] The term “control device” may generally refer to a network device configured to perform functions related to facilitating user access, control, and / or configuration of media playback system 100 .

[0016] Each of the playback devices 110 is configured to receive audio signals or data from one or more media sources (e.g., one or more remote servers or one or more local devices) and play the received audio signals or data as sound. One or more NMDs 120 are configured to receive voice word commands, and one or more control devices 130 are configured to receive user input. In response to the received spoken word commands and / or user input, the media playback system 100 can play audio via one or more of the playback devices 110. In certain examples, the playback devices 110 are configured to initiate playback of media content in response to a trigger. For example, one or more of the playback devices 110 can be configured to play a morning playlist upon detection of an associated trigger condition (e.g., a user's presence in the kitchen, detection of operation of the coffee machine). In some examples, for example, the media playback system 100 is configured to play audio from a first playback device (e.g., playback device 110a) in synchronization with a second playback device (e.g., playback device 110b). Interactions between playback device 110, NMD 120, and / or control device 130 of media playback system 100 configured according to various examples of the present disclosure are described in more detail below in conjunction with FIGS. 1B-1H.

[0017] 1A , environment 101 comprises a home with several rooms, spaces, and / or playback zones, including (clockwise from top left) master bathroom 101a, master bedroom 101b, second bedroom 101c, family room or den 101d, office 101e, living room 101f, dining room 101g, kitchen 101h, and outdoor patio 101i. While specific examples and examples are described below in the context of a home environment, the techniques described herein may be implemented in other types of environments. In some examples, for example, media playback system 100 may be implemented in one or more commercial settings (e.g., a restaurant, mall, airport, hotel, retail store, or other establishment), one or more vehicles (e.g., a sport utility vehicle, bus, car, watercraft, boat, airplane), multiple environments (e.g., a combination of home and vehicle environments), and / or other suitable environments where multi-zone audio may be desirable.

[0018] Media playback system 100 can include one or more playback zones, some of which may correspond to rooms within environment 101. Media playback system 100 can be established with one or more playback zones, after which additional zones can be added or removed to form the configuration shown in FIG. 1A, for example. Each zone can be named according to a different room or space, such as office 101e, master bathroom 101a, master bedroom 101b, second bedroom 101c, kitchen 101h, dining room 101g, living room 101f, and / or outdoor patio 101i. In some aspects, a single playback zone can include multiple rooms or spaces. In certain aspects, a single room or space can include multiple playback zones.

[0019] In the illustrated example of FIG. 1A , the master bathroom 101a, the second bedroom 101c, the office 101e, the living room 101f, the dining room 101g, the kitchen 101h, and the outdoor patio 101i each include one playback device 110, while the master bedroom 101b and the private room 101d include multiple playback devices 110. In the master bedroom 101b, the playback devices 110l and 110m may be configured to play audio content synchronously, for example, as individual ones of the playback devices 110, as a combined playback zone, as an integrated playback device, and / or any combination thereof. Similarly, in the private room 101d, the playback devices 110h-110j may be configured to play audio content synchronously, for example, as individual ones of the playback devices 110, as one or more combined playback devices, and / or as one or more integrated playback devices. Further details regarding combined and integrated playback devices are described below with respect to FIGS. 1B and 1E .

[0020] In some aspects, one or more of the playback zones in environment 101 may each be playing different audio content. For example, a user may be grilling on patio 101i and listening to hip hop music being played by playback device 110c, while another user may be preparing food in kitchen 101h and listening to classical music being played by playback device 110b. In another example, a playback zone may play the same audio content in sync with another playback zone. For example, a user may be in office 101e listening to playback device 110f playing the same hip hop music being played by playback device 110c on patio 101i. In some aspects, playback devices 110c and 110f play the hip hop music in sync so that the user perceives that the audio content is being played seamlessly (or at least substantially seamlessly) while moving between the different playback zones. Further details regarding audio playback synchronization between playback devices and / or zones can be found, for example, in U.S. Pat. No. 8,234,395, entitled "System and Method for Synchronizing Operation Among Multiple Independently Clocked Digital Data Processing Devices," which is incorporated herein by reference in its entirety.

[0021] a. Suitable media playback system 1B is a schematic diagram of media playback system 100 and cloud network 102. For ease of illustration, certain devices of media playback system 100 and cloud network 102 have been omitted from FIG. 1B. One or more communication links 103 (hereinafter "links 103") communicatively couple media playback system 100 and cloud network 102.

[0022] Link 103 may comprise, for example, one or more wired networks, one or more wireless networks, one or more wide area networks (WANs), one or more local area networks (LANs), one or more personal area networks (PANs), one or more telecommunications networks (e.g., one or more Global System for Mobile Communications (GSM) networks, Code Division Multiple Access (CDMA) networks, Long Term Evolution (LTE) networks, 5G communications networks, and / or other suitable data transmission protocol networks), etc. Cloud network 102 is configured to deliver media content (e.g., audio content, video content, photos, social media content) to media playback system 100 in response to requests transmitted from media playback system 100 via link 103. In some examples, cloud network 102 is further configured to receive data (e.g., voice input data) from media playback system 100 and, in response, transmit commands and / or media content to media playback system 100.

[0023] Cloud network 102 comprises computing devices 106 (separately identified as first computing device 106a, second computing device 106b, and third computing device 106c). Computing devices 106 may comprise individual computers or servers, such as, for example, media streaming service servers that store audio and / or other media content, voice service servers, social media servers, media playback system control servers, etc. In some examples, one or more of computing devices 106 comprise modules of a single computer or server. In particular examples, one or more of computing devices 106 comprise one or more modules, computers, and / or servers. Furthermore, while cloud network 102 is described above in the context of a single cloud network, in some examples, cloud network 102 comprises multiple cloud networks including communicatively coupled computing devices. Furthermore, while cloud network 102 is illustrated in FIG. 1B as having three of computing devices 106, in some examples, cloud network 102 comprises fewer (or more) than three computing devices 106.

[0024] Media playback system 100 is configured to receive media content from network 102 via link 103. The received media content may comprise, for example, a uniform resource identifier (URI) and / or a uniform resource locator (URL). For example, in some examples, media playback system 100 may stream, download, or retrieve data from a URI or URL corresponding to the received media content. Network 104 communicatively couples link 103 to at least some of the devices of media playback system 100 (e.g., one or more of playback device 110, NMD 120, and / or control device 130). Network 104 may include, for example, a wireless network (e.g., a WiFi network, Bluetooth, Z-Wave network, ZigBee, and / or other suitable wireless communication protocol network) and / or a wired network (e.g., a network including Ethernet, Universal Serial Bus (USB), and / or another suitable wired communication protocol). As will be appreciated by those skilled in the art, as used herein, "WiFi" can refer to several different communication protocols including, for example, Institute of Electrical and Electronics Engineers (IEEE) 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.11ac, 802.11ad, 802.11af, 802.11ah, 802.11ai, 802.11aj, 802.11aq, 802.11ax, 802.11ay, 802.15, etc., transmitted at 2.4 gigahertz (GHz), 5 GHz, and / or other suitable frequencies.

[0025] In some examples, network 104 comprises a dedicated communications network that media playback system 100 uses to send messages between individual devices and / or to transmit media content to and from media content sources (e.g., one or more of computing devices 106). In particular examples, network 104 is configured to be accessible only to devices within media playback system 100, thereby reducing interference and contention with other home devices. However, in other examples, network 104 comprises an existing home communications network (e.g., a home WiFi network). In some examples, link 103 and network 104 comprise one or more of the same networks. In some aspects, for example, link 103 and network 104 comprise a telecommunications network (e.g., an LTE network, a 5G network). Furthermore, in some examples, media playback system 100 is implemented without network 104, and devices comprising media playback system 100 can communicate with each other via, for example, one or more direct connections, PANs, telecommunications networks, and / or other suitable communications links.

[0026] In some examples, audio content sources may be periodically added or removed from media playback system 100. In some examples, for example, media playback system 100 performs media item indexing when one or more media content sources are updated, added, and / or removed from media playback system 100. Media playback system 100 may scan for identifiable media items in some or all folders and / or directories accessible to playback device 110 and generate or update a media content database comprising metadata (e.g., title, artist, album, track length) and other associated information (e.g., URI, URL) for each identifiable media item found. In some examples, for example, the media content database is stored on one or more of playback device 110, NMD 120, and / or control device 130.

[0027] In the illustrated example of FIG. 1B , playback devices 110l and 110m comprise group 107a. Playback devices 110l and 110m may be located in different rooms within a home and may be temporarily or permanently grouped together in group 107a based on user input received at control device 130a and / or another control device 130 within media playback system 100. Once arranged in group 107a, playback devices 110l and 110m may be configured to synchronously play the same or similar audio content from one or more audio content sources. In certain examples, for example, group 107a may include a combined zone in which playback devices 110l and 110m each include a left audio channel and a right audio channel of multi-channel audio content, thereby creating or enhancing a stereo effect of the audio content. In some examples, group 107a may include additional playback devices 110. However, in other examples, media playback system 100 omits group 107a and / or other grouped arrangements of playback devices 110.

[0028] Media playback system 100 includes NMDs 120a and 120d, each comprising one or more microphones configured to receive voice utterances from a user. In the illustrated example of FIG. 1B , NMD 120a is a standalone device, and NMD 120d is incorporated into playback device 110n. NMD 120a is configured to receive voice input 121, for example, from user 123. In some examples, NMD 120a transmits data related to the received voice input 121 to a voice assistant service (VAS) configured to (i) process the received voice input data and (ii) transmit corresponding commands to media playback system 100. In some aspects, for example, computing device 106c comprises one or more modules and / or servers of a VAS (e.g., SONOS®, AMAZON®, GOOGLE®, APPLE®, MICROSOFT®). Computing device 106c can receive the voice input data from NMD 120a via network 104 and link 103. In response to receiving the audio input data, computing device 106c processes the audio input data (i.e., "Play Hey Jude by the Beatles") and determines that the processed audio input includes a command to play a song (e.g., "Hey Jude"). Accordingly, computing device 106c transmits a command to media playback system 100 from one or more appropriate media services of playback devices 110 (e.g., via one or more of computing devices 106).

[0029] b. Suitable playback device 1C is a block diagram of a playback device 110a including an input / output 111. The input / output 111 can include an analog I / O 111a (e.g., one or more wires, cables, and / or other suitable communication links configured to carry analog signals) and / or a digital I / O 111b (e.g., one or more wires, cables, or other suitable communication links configured to carry digital signals). In some examples, the analog I / O 111a is an audio line-in connection, including, for example, an auto-sensing 3.5mm audio line-in connection. In some examples, the digital I / O 111b includes a Sony / Philips Digital Interface Format (S / PDIF) communication interface and / or cable and / or a Toshiba Link (TOSLINK) cable. In some examples, the digital I / O 111b includes a High-Definition Multimedia Interface (HDMI®) interface and / or cable. In some examples, Digital I / O 111b includes one or more wireless communication links, including, for example, radio frequency (RF), infrared, WiFi, Bluetooth, or another suitable communication protocol. In particular examples, Analog I / O 111a and Digital 111b include interfaces (e.g., ports, plugs, jacks) configured to accept connectors of cables that transmit analog and digital signals, respectively, without necessarily including cables.

[0030] Playback device 110a can receive media content (e.g., audio content including music and / or other sounds) from local audio source 105, for example, via input / output 111 (e.g., cable, wire, PAN, Bluetooth connection, ad-hoc wired or wireless communication network, and / or another suitable communication link). Local audio source 105 can comprise, for example, a mobile device (e.g., a smartphone, a tablet, a laptop computer) or another suitable audio component (e.g., a television, a desktop computer, an amplifier, a phonograph, a Blu-ray player, memory for storing digital media files). In some aspects, local audio source 105 includes a local music library on a smartphone, a computer, a network-attached storage (NAS), and / or another suitable device configured to store media files. In certain examples, one or more of playback device 110, NMD 120, and / or control device 130 comprise local audio source 105. However, in other examples, the media playback system omits local audio source 105 entirely. In some examples, playback device 110 a does not include input / output 111 and receives all audio content over network 104 .

[0031] Playback device 110a further comprises electronics 112, a user interface 113 (e.g., one or more buttons, knobs, dials, touch-sensitive surfaces, displays, touchscreens), and one or more transducers 114 (hereinafter referred to as “transducers 114”). Electronics 112 is configured to receive audio from an audio source (e.g., local audio source 105) via input / output 111, one or more of computing devices 106a-106c via network 104 (FIG. 1B), amplify the received audio, and output the amplified audio for playback via one or more of transducers 114. In some examples, playback device 110a optionally includes one or more microphones 115 (e.g., a single microphone, multiple microphones, a microphone array) (hereinafter referred to as “microphones 115”). In particular examples, for example, playback device 110a having one or more of optional microphones 115 may operate as an NMD configured to receive audio input from a user and correspondingly perform one or more actions based on the received audio input.

[0032] 1C , electronic device 112 includes one or more processors 112a (hereinafter referred to as “processor 112a”), memory 112b, software components 112c, network interface 112d, one or more audio processing components 112g (hereinafter referred to as “audio components 112g”), one or more audio amplifiers 112h (hereinafter referred to as “amplifiers 112h”), and power source 112i (e.g., one or more power supplies, power cables, power receptacles, batteries, induction coils, Power over Ethernet (POE) interfaces, and / or other suitable power sources). In some embodiments, electronic device 112 optionally includes one or more other components 112j (e.g., one or more sensors, a video display, a touch screen, a battery charging base).

[0033] The processor 112a may comprise a clocked computing component configured to process data, and the memory 112b may comprise a computer-readable medium (e.g., a tangible, non-transitory computer-readable medium, data storage device) configured to store instructions for performing various operations and / or functions. The processor 112a is configured to execute the instructions stored in the memory 112b to perform one or more of the operations. The operations may include, for example, causing the playback device 110a to retrieve audio data from an audio source (e.g., one or more of the computing devices 106a-106c (FIG. 1B)) and / or another one of the playback devices 110. In some examples, the operations further include causing the playback device 110a to transmit the audio data to another one of the playback devices 110a and / or to another device (e.g., one of the NMDs 120). Particular examples include pairing the playback device 110a with another of the one or more playback devices 110 to enable a multi-channel audio environment (e.g., stereo pair, combined zone).

[0034] The processor 112a may be further configured to perform operations that cause the playback device 110a to synchronize playback of the audio content with another of the one or more playback devices 110. As will be appreciated by those skilled in the art, during synchronized playback of audio content on multiple playback devices, a listener preferably cannot perceive a time delay difference between the playback of the audio content by the playback device 110a and the playback of the audio content by one or more other playback devices 110. Further details regarding audio playback synchronization between playback devices may be found, for example, in U.S. Patent No. 8,234,395, incorporated by reference above.

[0035] In some examples, memory 112b is further configured to store data associated with playback device 110a, such as one or more zones and / or zone groups of which playback device 110a is a member, audio sources accessible to playback device 110a, and / or playback queues to which playback device 110a (and / or other playback devices of the one or more playback devices) may be associated. The stored data may include one or more state variables that are periodically updated and used to describe the state of playback device 110a. Memory 112b may also include data associated with the state of one or more of the other devices of media playback system 100 (e.g., playback device 110, NMD 120, control device 130). In some aspects, for example, state data is shared among at least some of the devices of media playback system 100 at predetermined time intervals (e.g., every 5 seconds, every 10 seconds, every 60 seconds), so that one or more of the devices have up-to-date data associated with media playback system 100.

[0036] Network interface 112d is configured to facilitate the transmission of data between playback device 110a and one or more other devices on a data network, such as link 103 and / or network 104 (FIG. 1B). Network interface 112d is configured to send and receive data corresponding to media content (e.g., audio content, video content, text, photographs) and other signals (e.g., non-transient signals) including digital packet data that include an Internet Protocol (IP)-based source address and / or an IP-based destination address. Network interface 112d can parse the digital packet data so that electronic device 112 appropriately receives and processes the data destined for playback device 110a.

[0037] In the depicted example of FIG. 1C , the network interface 112d includes one or more wireless interfaces 112e (hereinafter referred to as “wireless interface 112e”). The wireless interface 112e (e.g., a suitable interface including one or more antennas) can be configured to wirelessly communicate with one or more other devices (e.g., one or more of the other playback devices 110, the NMD 120, and / or the control device 130) communicatively coupled to the network 104 ( FIG. 1B ) according to a suitable wireless communication protocol (e.g., WiFi, Bluetooth, LTE). In some examples, the network interface 112d optionally includes a wired interface 112f (e.g., an interface or receptacle configured to receive a network cable, such as an Ethernet, USB-A, USB-C, and / or Thunderbolt cable) configured to communicate with other devices via a wired connection according to a suitable wired communication protocol. In particular examples, the network interface 112d includes the wired interface 112f and excludes the wireless interface 112e. In some examples, the electronic device 112 omits the network interface 112d entirely and transmits and receives media content and / or other data via another communication path (eg, input / output 111).

[0038] Audio component 112g is configured to process and / or filter data including media content received by electronic device 112 (e.g., via input / output 111 and / or network interface 112d) to generate an output audio signal. In some examples, audio processing component 112g comprises, for example, one or more digital-to-analog converters (DACs), audio pre-processing components, audio enhancement components, digital signal processors (DSPs), and / or other suitable audio processing components, modules, circuits, etc. In particular examples, one or more of audio processing components 112g may comprise one or more subcomponents of processor 112a. In some examples, electronic device 112 omits audio processing component 112g. In some aspects, for example, processor 112a executes instructions stored in memory 112b to perform audio processing operations to generate an output audio signal.

[0039] The amplifiers 112h are configured to receive and amplify audio output signals generated by the audio processing component 112g and / or the processor 112a. The amplifiers 112h may comprise electronic devices and / or components configured to amplify the audio signals to a level sufficient to drive one or more of the transducers 114. In some examples, for example, the amplifiers 112h include one or more switching or class-D power amplifiers. However, in other examples, the amplifiers include one or more other types of power amplifiers (e.g., linear gain power amplifiers, class-A amplifiers, class-B amplifiers, class-AB amplifiers, class-C amplifiers, class-D amplifiers, class-E amplifiers, class-F amplifiers, class-G amplifiers, and / or class-H amplifiers, and / or other suitable types of power amplifiers). In particular examples, the amplifiers 112h comprise a suitable combination of two or more of the aforementioned types of power amplifiers. Furthermore, in some examples, individual ones of the amplifiers 112h correspond to individual ones of the transducers 114. However, in other examples, the electronics 112 includes a single one of the amplifiers 112h configured to output an amplified audio signal to the plurality of transducers 114. In some other examples, the electronics 112 omits the amplifier 112h.

[0040] The transducer 114 (e.g., one or more speakers and / or speaker drivers) receives the amplified audio signal from the amplifier 112h and renders or outputs the amplified audio signal as sound (e.g., audible sound waves having a frequency between approximately 20 Hertz (Hz) and 20 Kilohertz (kHz)). In some examples, the transducer 114 may comprise a single transducer. However, in other examples, the transducer 114 comprises multiple audio transducers. In some examples, the transducer 114 comprises multiple types of transducers. For example, the transducer 114 may include one or more low-frequency transducers (e.g., subwoofers, woofers), a mid-frequency transducer (e.g., mid-range transducer, mid-woofer), and one or more high-frequency transducers (e.g., one or more tweeters). As used herein, "low frequency" can generally refer to audible frequencies below about 500 Hz, "mid-range frequency" can generally refer to audible frequencies between about 500 Hz and about 2 kHz, and "high frequency" can generally refer to audible frequencies above 2 kHz. However, in certain examples, one or more of the transducers 114 comprises a transducer that does not adhere to the aforementioned frequency ranges. For example, one of the transducers 114 may comprise a mid-woofer transducer configured to output sound at frequencies between about 200 Hz and about 5 kHz.

[0041] By way of example, SONOS, Inc. currently offers (or has offered) certain playback devices for sale, including, for example, “SONOS ONE,” “PLAY:1,” “PLAY:3,” “PLAY:5,” “PLAYBAR,” “PLAYBASE,” “CONNECT:AMP,” “CONNECT,” and “SUB.” Additionally or alternatively, other suitable playback devices may be used to implement the example playback devices disclosed herein. Furthermore, as will be appreciated by those skilled in the art, playback devices are not limited to the examples described herein or to the SONOS product offerings. In some examples, for example, one or more of the playback devices 110 comprise wired or wireless headphones (e.g., over-ear headphones, earbuds, in-ear earphones). In other examples, one or more of the playback devices 110 comprise a docking station and / or interface configured to interact with a docking station for a personal mobile media playback device. In certain examples, the playback device may be integrated with another device or component, such as a television, a lighting fixture, or some other device for indoor or outdoor use. In some examples, the playback device omits a user interface and / or one or more transducers. For example, FIG. 1D is a block diagram of a playback device 110 p with input / output 111 and electronics 112 without a user interface 113 or transducer 114 .

[0042] FIG. 1E is a block diagram of a combined playback device 110q comprising playback device 110i (e.g., a subwoofer) ( FIG. 1A ) and ultrasonically bonded playback device 110a ( FIG. 1C ). In the illustrated example, playback devices 110a and 110i are separate playback devices 110 housed in separate housings. However, in some examples, combined playback device 110q comprises a single housing housing both playback devices 110a and 110i. Combined playback device 110q can be configured to process and reproduce sound differently than uncoupled playback devices (e.g., playback device 110a of FIG. 1C ) and / or paired or combined playback devices (e.g., playback devices 110l and 110m of FIG. 1B ). In some examples, for example, playback device 110a is a full-range playback device configured to render low-, mid-, and high-frequency audio content, and playback device 110i is a subwoofer configured to render low-frequency audio content. In some aspects, playback device 110a, when coupled with a first playback device, is configured to render only the mid- and high-frequency components of a particular audio content, while playback device 110i renders the low-frequency components of the particular audio content. In some examples, combined playback device 110q includes additional playback devices and / or another combined playback device.

[0043] c. A suitable Network Microphone Device (NMD) FIG. 1F is a block diagram of NMD 120a (FIGS. 1A and 1B). NMD 120a includes one or more audio processing components 124 (hereinafter “audio components 124”) and several components described with respect to playback device 110a (FIG. 1C), including processor 112a, memory 112b, and microphone 115. NMD 120a optionally includes other components also included in playback device 110a (FIG. 1C), such as user interface 113 and / or transducer 114. In some examples, NMD 120a is configured as a media playback device (e.g., one or more of playback devices 110) and further includes, for example, one or more of audio component 112g (FIG. 1C), amplifier 114, and / or other playback device components. In particular examples, NMD 120a comprises an Internet of Things (IoT) device, such as a thermostat, an alarm panel, a fire and / or a smoke detector, etc. In some examples, NMD 120a includes microphone 115, audio processing 124, and only some of the components of electronics 112 described above with respect to FIG. 1B. In some aspects, for example, NMD 120a includes processor 112a and memory 112b (FIG. 1B) while omitting one or more other components of electronics 112. In some examples, NMD 120a includes additional components (e.g., one or more sensors, a camera, a thermometer, a barometer, a hygrometer).

[0044] In some examples, an NMD can be incorporated into a playback device. FIG. 1G is a block diagram of playback device 110r including NMD 120d. Playback device 110r can include many or all of the components of playback device 110a and can further include microphone 115 and audio processing 124 (FIG. 1F). Playback device 110r may have an integrated control device 130c. Control device 130c can include, for example, a user interface (e.g., user interface 113 of FIG. 1B) configured to receive user input (e.g., touch input, voice input) without a separate control device. However, in other examples, playback device 110r receives commands from another control device (e.g., control device 130a of FIG. 1B).

[0045] Referring again to FIG. 1F , the microphone 115 is configured to acquire, capture, and / or receive sound from the environment (e.g., environment 101 of FIG. 1A ) and / or room in which the NMD 120a is located. The received sound may include, for example, voice utterances, audio played by the NMD 120a and / or another playback device, background sounds, ambient sounds, etc. The microphone 115 converts the received sound into electrical signals to generate microphone data. The audio processing 124 receives and analyzes the microphone data to determine whether the microphone data contains voice input. The voice input may include, for example, a wake-up word followed by an utterance containing a user request. As will be appreciated by those skilled in the art, a wake-up word is a word or other audio cue signifying user voice input. For example, when querying an AMAZON® VAS, a user may utter the wake-up word “Alexa.” Other examples include “Okay, Google” to invoke a Google® VAS and “Hey, Siri” to invoke an Apple® VAS.

[0046] After detecting the hotword, voice processing 124 monitors microphone data for an accompanying user request in the voice input. The user request may include, for example, a command to control a third-party device such as a thermostat (e.g., a NEST® thermostat), a lighting device (e.g., a PHILIPS HUE® lighting device), or a media playback device (e.g., a Sonos® playback device). For example, a user may utter the hotword “Alexa” (e.g., environment 101 of FIG. 1A ) followed by the utterance “Set the thermostat to 68 degrees” to set the temperature in their home. A user may utter the same hotword followed by the utterance “Turn on the living room” to turn on a lighting device in the living room area of ​​their home. A user may similarly speak the hotword followed by a request to play a particular song, album, or music playlist on a playback device in their home.

[0047] d. Appropriate control devices FIG. 1H is a partial schematic diagram of control device 130a (FIGS. 1A and 1B). As used herein, the term "control device" can be used interchangeably with "controller" or "control system." Among other features, control device 130a is configured to receive user input associated with media playback system 100 and, in response, cause one or more devices within media playback system 100 to perform an action or actions corresponding to the user input. In the illustrated example, control device 130a comprises a smartphone (e.g., iPhone®, Android phone) having media playback system controller application software installed. In some examples, control device 130a comprises, for example, a tablet (e.g., iPad®), a computer (e.g., laptop computer, desktop computer), and / or another suitable device (e.g., television, automobile audio head unit, IoT device). In particular examples, control device 130a comprises a dedicated controller for media playback system 100. In other examples, as described above with respect to FIG. 1G, control device 130a is incorporated into another device in media playback system 100 (e.g., one or more of playback device 110, NMD 120, and / or other suitable devices configured to communicate over a network).

[0048] Control device 130a includes electronics 132, a user interface 133, one or more speakers 134, and one or more microphones 135. Electronics 132 includes one or more processors 132a (hereinafter referred to as “processor 132a”), memory 132b, software components 132c, and a network interface 132d. Processor 132a can be configured to perform functions related to facilitating user access, control, and configuration of media playback system 100. Memory 132b can include data storage into which one or more of the software components executable by processor 112a to perform these functions can be loaded. Software component 132c can include applications and / or other executable software configured to facilitate control of media playback system 100. Memory 112b can be configured to store, for example, software component 132c, media playback system controller application software, and / or other data related to media playback system 100 and users.

[0049] Network interface 132d is configured to facilitate network communication between control device 130a and one or more other devices in media playback system 100 and / or one or more remote devices. In some examples, network interface 132d is configured to operate according to one or more appropriate communications industry standards (e.g., infrared, wireless, wired standards including IEEE 802.3, wireless standards including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G, LTE). Network interface 132d can be configured to transmit and / or receive data to, for example, playback device 110, NMD 120, other control devices 130, one of computing devices 106 of FIG. 1B, one or more other devices comprising the media playback system, etc. Transmitted and / or received data can include, for example, playback device control commands, state variables, playback zones and / or zone group configurations. For example, based on user input received at user interface 133, network interface 132d can transmit playback device control commands (e.g., volume control, audio playback control, audio content selection) from control device 130 to one or more of playback devices 110. Network interface 132d can also transmit and / or receive configuration changes such as, for example, adding / removing one or more playback devices 110 to / from a zone, adding / removing one or more zones to / from a zone group, forming a combined or integrated player, separating one or more playback devices from a combined or integrated player, among others. A description of adding zones and groups can be found below with respect to Figures 1I-1M.

[0050] The user interface 133 is configured to receive user input and can facilitate control of the media playback system 100. The user interface 133 includes media content techniques 133a (e.g., album art, lyrics, video), playback status indicators 133b (e.g., elapsed time and / or time remaining indicators), media content information area 133c, playback control area 133d, and zone indicators 133e. The media content information area 133c can include a display of relevant information (e.g., title, artist, album, genre, release year) about the currently playing media content and / or media content in a queue or playlist. The playback control area 133d can include selectable (e.g., via touch input and / or via a cursor or another suitable selector) icons to cause one or more playback devices in a selected playback zone or zone group to perform playback actions, such as play or pause, fast forward, rewind, skip next, skip previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit crossfade mode, etc. Playback control area 133d may also include selectable icons for changing equalization settings, playback volume, and / or other appropriate playback operations. In the illustrated example, user interface 133 comprises a display presented on a touchscreen interface of a smartphone (e.g., iPhone®, Android phone). However, in some examples, user interfaces of various formats, styles, and interactive sequences may alternatively be implemented on one or more networked devices to provide equivalent control access to a media playback system.

[0051] One or more speakers 134 (e.g., one or more transducers) may be configured to output audio to a user of control device 130a. In some examples, one or more speakers comprise individual transducers configured to output corresponding low, mid, and / or high frequencies. In some aspects, for example, control device 130a is configured as a playback device (e.g., one of playback devices 110). Similarly, in some examples, control device 130a is configured as an NMD (e.g., one of NMDs 120) and receives voice commands and other audio via one or more microphones 135.

[0052] The one or more microphones 135 may comprise, for example, one or more condenser microphones, electret condenser microphones, dynamic microphones, and / or other suitable types of microphones or transducers. In some examples, two or more of the microphones 135 are positioned to capture location information of an audio source (e.g., voice, audible sound) and / or configured to facilitate filtering of background noise. Furthermore, in certain examples, the control device 130a is configured to operate as a playback device and an NMD. However, in other examples, the control device 130a omits one or more speakers 134 and / or one or more microphones 135. For example, the control device 130a may comprise a device (e.g., a thermostat, an IoT device, a network device) that includes a portion of the electronics 132 and a user interface 133 (e.g., a touchscreen) without a speaker or microphone.

[0053] Proper playback device configuration 1I-1M show exemplary configurations of playback devices in zones and zone groups. Referring first to FIG. 1M, in one example, a single playback device can belong to a zone. For example, playback device 110g in the second bedroom 101c (FIG. 1A) may belong to Zone C. In some implementations described below, multiple playback devices can be "combined" to form a "combined pair," which together form a single zone. For example, playback device 110l (e.g., the left playback device) can be combined with playback device 110j (e.g., the right playback device) to form Zone A. The combined playback devices may have different playback responsibilities (e.g., channel responsibilities). In another implementation described below, multiple playback devices can be merged to form a single zone. For example, playback device 110h (e.g., the front playback device) can be merged with playback device 110i (e.g., a subwoofer) and playback devices 110j and 110k (e.g., left and right surround speakers, respectively) to form a single Zone D. In another example, playback devices 110g and 110h can be merged to form merged group or zone group 108b. Merged playback devices 110g and 110h may not be specifically assigned different playback responsibilities. That is, merged playback devices 110h and 110i can play audio content the same way as if they were not merged, apart from playing audio content synchronously.

[0054] Each zone within media playback system 100 may be provided for control as a single user interface (UI) entity. For example, Zone A may be provided as a single entity called the master bathroom. Zone B may be provided as a single entity called the master bedroom. Zone C may be provided as a single entity called the second bedroom.

[0055] Combined playback devices may have different playback responsibilities, such as responsibility for specific audio channels. For example, as shown in FIG. 1-I, playback devices 110l and 110m may be combined to create or enhance a stereo effect for audio content. In this example, playback device 110l may be configured to play the left channel audio component, while playback device 110k may be configured to play the right channel audio component. In some implementations, such stereo combining may be referred to as "pairing."

[0056] Furthermore, combined playback devices may have additional and / or different respective speaker drivers. As shown in FIG. 1J, a playback device 110h labeled Front may be combined with a playback device 110i labeled SUB. The Front device 110h may be configured to render a mid- to high-frequency range, and the SUB device 110i may be configured to render low frequencies. However, when uncombined, the Front device 110h may be configured to render the entire frequency range. As another example, FIG. 1K shows the Front device 110h and the SUB device 110i further combined with the left playback device 110j and the right playback device 110k, respectively. In some implementations, the right device 110j and the left device 110k may be configured to form a surround or "satellite" channel in a home theater system. The combined playback devices 110h, 110i, 110j, and 110k may form a single zone D (FIG. 1M).

[0057] Merged playback devices may not have assigned playback responsibilities and may be capable of rendering the full range of audio content for which each playback device is capable. Nevertheless, merged devices may be represented as a single UI entity (i.e., a zone, as described above). For example, playback devices 110a and 110n in the master bathroom have a single UI entity for Zone A. In one example, playback devices 110a and 110n can each output the full range of audio content for which each playback device 110a and 110n is capable in sync.

[0058] In some examples, an NMD is combined or merged with another device to form a zone. For example, NMD 120b may be combined with playback device 110e, which together form zone F, referred to as the living room. In other examples, a standalone network microphone device may itself be within a zone. However, in other examples, a standalone network microphone device may not be associated with a zone. Further details regarding associating network microphone devices and playback devices as designated or default devices can be found, for example, in the above-referenced U.S. patent application Ser. No. 15 / 438,749.

[0059] Zones of individual devices, combined devices, and / or merged devices may be grouped to form zone groups. For example, referring to FIG. 1M, zone A may be grouped with zone B to form zone group 108a containing the two zones. Similarly, zone G may be grouped with zone H to form zone group 108b. As another example, zone A may be grouped with one or more other zones C1. Zones A-I may be grouped and ungrouped in numerous ways. For example, three, four, five, or more (e.g., all) of zones A-I may be grouped. Once grouped, zones of individual and / or combined playback devices may play audio in synchronization with one another, as described in the previously referenced U.S. Patent No. 8,234,395. Playback devices may also be dynamically grouped and ungrouped to form new or different groups that play audio content in synchronization.

[0060] In various implementations, a zone within an environment may be a combination of the default names of the zones within the group or the names of the zones within the zone group. For example, zone group 108b may be assigned a name such as "dining + kitchen," as shown in FIG. 1M. In some examples, a zone group may be given a unique name selected by the user.

[0061] Certain data may be stored in the memory of the playback device (e.g., memory 112b of FIG. 1C) as one or more state variables that are periodically updated and used to describe the state of the playback zone, the playback device, and / or the zone group associated with the playback zone. The memory may also contain data that is associated with the state of other devices in the media system and that is shared from time to time between devices so that one or more of the devices have the most current data associated with the system.

[0062] In some examples, the memory may store instances of various variable types associated with states. The variable instances may be stored with an identifier (e.g., a tag) corresponding to the type. For example, a particular identifier may be a first type "a1" to identify a playback device in a zone, a second type "b1" to identify playback devices that can be combined within the zone, and a third type "c1" to identify a zone group to which the zone can belong. As a related example, an identifier associated with the second bedroom 101c may indicate that the playback device is the only playback device in zone C and not within a zone group. An identifier associated with Den may indicate that Den is not grouped with other zones but includes combined playback devices 110h-110k. An identifier associated with the dining room may indicate that the dining room is part of the dining + kitchen zone group 108b and that devices 110b and 110d are grouped together (FIG. 1L). An identifier associated with the kitchen may indicate the same or similar information by virtue of the kitchen being part of the dining + kitchen zone group 108b. Other exemplary zone variables and identifiers are described below.

[0063] In yet another example, media playback system 100 may include variables or identifiers representing other associations of zones and zone groups, such as identifiers associated with areas, as shown in FIG. 1M. Areas can include clusters of zone groups and / or zones not within a zone group. For example, FIG. 1M shows upper area 109a including zones A-D and lower area 109b including zones E-I. In one aspect, an area can be used to refer to a zone group and / or cluster of zones that share one or more zones and / or zone groups of another cluster. In another aspect, this is distinct from a zone group that does not share a zone with another zone group. Further examples of techniques for implementing areas can be found, for example, in U.S. Patent Application No. 15 / 682,506, filed August 21, 2017, entitled "Name-Based Room Association," and U.S. Patent No. 8,483,853, filed September 11, 2007, entitled "Control and Manipulation of Groupings in a Multi-Zone Media System." Each of these applications is incorporated herein by reference in its entirety. In some examples, media playback system 100 may not implement areas, in which case the system may not store variables associated with areas.

[0064] III. Multi-device playback of generated media content FIG. 2 is a functional block diagram of a system 200 for the playback of generated media content. As previously mentioned, generated media content can include any media content (e.g., audio, video, audiovisual output, tactile output, or any other media content) that is dynamically created, synthesized, and / or modified by a non-human, rule-based process, such as an algorithm or model. This creation or modification can occur for playback in real time or near real time. Additionally or alternatively, generated media content can be generated or modified asynchronously (e.g., before playback is requested), and specific items of generated media content can then be selected for playback at a later time. As used herein, a "generative media module" includes any system capable of generating generated media content based on one or more inputs, whether implemented with software, a physical model, or a combination thereof. In some examples, such generated media content includes new media content that can be created entirely new or by mixing, combining, manipulating, or otherwise modifying one or more existing pieces of media content. As used herein, a "generated media content model" includes any algorithm, schema, or set of rules that can be used to generate new generated media content using one or more inputs (e.g., sensor data, artist-provided parameters, media segments such as audio clips or samples, etc.). In examples, a generative media module can use a variety of different generated media content models to generate different generated media content. In some cases, an artist or other collaborator can interact with, create, and / or update a generated media content model to generate particular generated media content.Although some examples throughout this description refer to audio content, the principles disclosed herein may, in some examples, be applied to other types of media content, such as video, audiovisual, tactile, or other.

[0065] 2, system 200 includes a produced media group coordinator 210 that communicates with produced media group members 250a and 250b, as well as a sensor data source 218, a media content source 220, and a control device 130. Such communication may be performed over network 102, which, as previously described, may include any suitable wired or wireless network connection or combination thereof (e.g., a WiFi network, Bluetooth, a Z-Wave network, ZigBee, an Ethernet connection, a Universal Serial Bus (USB) connection, etc.).

[0066] One or more remote computing devices 106 can also communicate with the group coordinator 210 and / or group members 250a and 250b via the network 102. In various examples, the remote computing device 106 can be a cloud-based server associated with a device manufacturer, a media content provider, a voice assistant service, or other suitable entity. As shown in FIG. 2, the remote computing device 106 can include a generated media module 214. As described in more detail elsewhere herein, the remote computing device 106 can generate generated media content remotely from local devices (e.g., the coordinator 210 and members 250a and 250b). The generated media content can then be transmitted to one or more local devices for playback. Additionally or alternatively, the generated media content can be generated in whole or in part via local devices (e.g., the group coordinator 210 and / or group members 250a and 250b). In some examples, the group coordinator 210 may itself be a remote computing device, communicatively coupled to group members 250a and 250b via a wide area network, and the devices need not be co-located within the same environment (e.g., home, business, etc.).

[0067] a. Example of generated media group behavior In the depicted example, the produced media group includes a produced media group coordinator 210 (also referred to herein as “coordinator device 210”) and first and second produced media group members 250 a and 250 b (also referred to herein collectively as “first member device 250 a,” “second member device 250 b,” and “member devices 250”). Optionally, one or more remote computing devices 106 may also form part of the produced media group. In operation, these devices may communicate with each other and / or other components (e.g., sensor data source 218, control device 130, media content source 220, or any other suitable data source or component) to facilitate the generation and playback of produced media content.

[0068] In various examples, some or all of devices 210 and / or 250 may be co-located in the same environment (e.g., in the same home, in a store, etc.) In some examples, at least some of devices 210 and / or 250 may be remote from one another, e.g., in different homes, in different cities, etc.

[0069] 1A-1H, the coordinator device 210 and / or the member device 250 may include some or all of the components of the playback device 110 or the network microphone device 120. For example, the coordinator device 210 and / or the member device 250 may optionally include playback components 212 (e.g., transducers, amplifiers, audio processing components, etc.), or such components may be omitted in some cases.

[0070] In some examples, the coordinator device 210 is a playback device itself and therefore can also operate as a member device 250. In other examples, the coordinator device 210 can be connected to one or more member devices 250 (e.g., via a direct wired connection or via network 102), but the coordinator device 210 does not itself play the produced media content. In various examples, the coordinator device 210 can be implemented on a bridge-like device on a local network, a playback device that is not itself part of a produced media group (i.e., the playback device does not itself play the produced media content), and / or a remote computing device (e.g., a cloud server).

[0071] In various examples, one or more of the devices may include a generated media module 214 thereon. Such generated media module 214 may generate new composite media content based on one or more inputs, for example, using an appropriate generated media content model. As shown in FIG. 2 , in some examples, the coordinator device 210 may include a generated media module 214 for generating generated media content, which may then be transmitted to member devices 250 a and 250 b for simultaneous and / or synchronized playback. Additionally or alternatively, some or all of the member devices 250 (e.g., member device 250 b shown in FIG. 2 ) may include a generated media module 214, which may be used by member device 250 to locally generate generated media content based on one or more inputs. In various examples, the generated media content may be generated via the remote computing device 106, optionally using one or more input parameters received from a local device. This generated media content may then be transmitted to one or more local devices for coordination and / or playback.

[0072] In some examples, at least some of the member devices 250 do not include a generated media module 214 therein. Alternatively, in some cases, each member device 250 can include a generated media module 214 therein and can be configured to generate generated media content locally. In at least some examples, none of the member devices 250 include a generated media module 214 therein. In such cases, generated media content can be generated by the coordinator device 210. Such generated media content can then be transmitted to the member devices 250 for simultaneous and / or synchronized playback.

[0073] 2, coordinator device 210 further includes coordination component 216. As described in more detail herein, in some cases, coordinator device 210 can facilitate playback of produced media content via multiple different playback devices (which may or may not include coordinator device 210 itself). In operation, coordination component 216 is configured to facilitate synchronization of both produced media creation (e.g., using one or more produced media modules 214, which may be distributed among various devices) and produced media playback. For example, coordinator device 210 can transmit timing data to member devices 250 to facilitate synchronized playback. Additionally or alternatively, the coordinator device 210 may send input, generated media model parameters, or other data regarding the generated media modules 214 to one or more member devices 250 so that the member devices 250 can generate the generated media locally (e.g., using locally stored generated media modules 214) and / or so that the member devices 250 can update or modify the generated media modules 214 based on input received from the coordinator device 210.

[0074] As described in more detail elsewhere herein, the generated media module 214 can be configured to generate generated media based on one or more inputs using a generated media content model. The inputs can include sensor data (e.g., as provided by a sensor data source 218), user input (e.g., as received from the control device 130 or via direct user interaction with the coordinator device 210 or a member device 250), and / or a media content source 220. For example, the generated media module 214 can generate and continuously modify the generated audio by adjusting various characteristics of the generated audio based on one or more input parameters (e.g., sensor data related to one or more users of the devices 210, 250).

[0075] b. Exemplary Media Content Sources Media content source 220, in various examples, can include one or more local and / or remote media content sources. For example, media content source 220 can include one or more local audio sources 105 (e.g., audio received via an input / output connection from a mobile device (e.g., a smartphone, tablet, laptop computer) or another suitable audio component (e.g., a television, desktop computer, amplifier, phonograph, Blu-ray player, memory storing digital media files), etc.) as described above. Additionally or alternatively, media content source 220 can include one or more remote computing devices accessible via a network interface (e.g., via communication over network 102). Such remote computing devices can include, for example, individual computers or servers, such as media streaming service servers, that store audio and / or other media content, etc.

[0076] In various examples, media available via media content source 220 may include pre-recorded audio segments in the form of complete sounds, songs, portions of songs (e.g., samples), or any audio component (e.g., pre-recorded audio of a particular instrument, synthesized beats or other audio segments, non-musical audio such as spoken word or natural sounds, etc.). In operation, such media may be utilized by generated media module 214 to generate generated media content, for example, by combining, mixing, overlaying, manipulating, or otherwise modifying the retrieved media content to generate new generated media content for playback via one or more devices. In some examples, the generated media content may take the form of a combination of pre-recorded audio segments (e.g., pre-recorded songs, spoken word recordings, etc.) and new synthesized audio that is created and overlaid with the pre-recorded audio. As used herein, “generated media content” or “generated media content” may include any such combination.

[0077] c. Exemplary Generative Media Module As previously mentioned, the generative media module 214 may include any system capable of generating generative media content based on one or more inputs, whether instantiated in software, a physical model, or a combination thereof. In various examples, the generative media module 214 may utilize a generative media content model, which may include one or more algorithms or mathematical models that determine how media content is generated based on associated input parameters. In some cases, the algorithms and / or mathematical models themselves may be updated over time, for example, based on instructions received from one or more remote computing devices (e.g., a cloud server associated with a music service or other entity), or based on input received from other group member devices in the same or different environments, or any other suitable input. In some examples, various devices in a group may have different generative media modules 214 thereon, for example, with a first member device having a different generative media module 214 than a second member device. In other cases, each device in a group having a generative media module 214 may include substantially the same model or algorithm.

[0078] Any suitable algorithm or combination of algorithms may be used to generate the generative media content. Examples of such algorithms include those using machine learning techniques (e.g., generative adversarial networks, neural networks, etc.), formal grammars, Markov models, finite-state automata, and / or any algorithms implemented in currently available offerings such as JukeBox by OpenAI, AWS DeepComposer by Amazon, Magenta by Google, and Amper AI by Amper Music. In various examples, the generative media module 214 may utilize any suitable generative algorithm currently existing or developed in the future.

[0079] In line with the above description, generating generative media content (e.g., audio content) can include modifying various characteristics of the media content in real time and / or algorithmically generating new media content in real time or near real time. In the context of audio content, this can be accomplished by storing several audio samples in a database (e.g., within the media content source 220), which can be remotely located and accessible by the coordinator device 210 and / or the member devices 250 via the network 102, or the audio samples can be maintained locally on the devices 210, 250 themselves. The audio samples can be associated with one or more metadata tags corresponding to one or more audio characteristics of the sample. For example, a given sample can be associated with metadata tags indicating that the sample contains audio of a particular frequency or frequency range (e.g., bass / midrange / treble), or a particular instrument, genre, tempo, key, release date, geographic region, timbre, reverb, distortion, sonic texture, or any other audio characteristic that becomes apparent.

[0080] In operation, the generative media module 214 (e.g., of the coordinator device 210 and / or the second member device 250b) can search for specific audio samples based on their associated tags and blend the audio samples to create the generated audio. The generated audio can evolve in real time as the generative media module 214 searches for audio samples with different tags and / or different audio samples with the same or similar tags. The audio samples that the generative media module(s) 214 search for can depend on one or more inputs, such as sensor data, time of day, geographic location, weather, or various user inputs, such as mood selection, or physiological inputs, such as heart rate. In this manner, as the inputs change, the generated audio also changes. For example, if a user selects a calming or relaxing mood input, the generative media module(s) 214 can search for and blend audio samples with tags corresponding to audio content that the user finds calming or relaxing. Examples of such audio samples can include audio samples tagged as low tempo or low harmonic complexity, or audio samples predetermined and tagged as calm and relaxing. In some examples, audio samples may be identified as calming or relaxing based on an automated process that analyzes the temporal and spectral content of the signal. Other examples are possible as well. In any of the examples herein, the generated media module 214 may adjust the characteristics of the generated audio by searching for and mixing audio samples associated with different metadata tags or other suitable identifiers.

[0081] Modifying the characteristics of the generated audio can include manipulating one or more of the volume, balance, removal of particular instruments or tones, changing the tempo, gain, reverb, spectral equalization, timbre, or sound texture of the audio, etc. In some examples, the generated audio can be played back differently on different devices, such as by emphasizing particular characteristics of the generated audio on a particular playback device closest to the user. For example, the closest playback device can emphasize a particular instrument, beat, tone, or other characteristic, while the remaining playback devices can serve as background audio sources.

[0082] As described elsewhere herein, the media content module 214 can be configured to generate media intended to steer the user's mood and / or physiological state in a desired direction. In some examples, the user's current state (e.g., mood, emotional state, activity level, etc.) is constantly and / or repeatedly monitored or measured (e.g., at predetermined intervals) to ensure that the user's current state is moving toward the desired state or at least not moving in the opposite direction to the desired state. In such examples, the generated audio content can be altered to move the user's current state toward a desired end state.

[0083] In any of the examples herein, the generative media module may use hysteresis to avoid rapid adjustments to the generated audio that could adversely affect the listening experience. For example, if the generative media module changes media based on user position input relative to the playback device, the playback device may rapidly change the generated audio in any of the methods described herein when the user moves closer to or further away from the playback device. Such abrupt adjustments may be unpleasant for the user. To reduce these rapid adjustments, the generative media module 214 may be configured to use hysteresis by delaying adjustments to the generated audio for a predetermined period of time when user movement or other activity triggers an adjustment. For example, if the playback device detects that the user has moved within a threshold distance of the playback device, instead of immediately performing one of the adjustments described above, the playback device may wait a predetermined amount of time (e.g., several seconds) before making the adjustment. If the user remains within the threshold distance after the predetermined amount of time, the playback device may proceed to adjust the generated audio. However, if the user does not remain within the threshold distance after the predetermined amount of time, the generative media module 214 may refrain from adjusting the generated audio. The generated media module 214 can similarly apply hysteresis to other generated media adjustments described herein.

[0084] FIG. 3 shows a flowchart of a process 300 for generating generated audio content using various input parameters. In various examples, one or more of these input parameters can be modified based on user input. For example, an artist can select various parameters, constraints, or available audio segments shown in FIG. 2, and these selections can at least partially determine the final output of the generated audio content. As previously described, such generated media modules may be stored and manipulated on one or more playback devices for local playback (e.g., via the same playback device and / or via other playback devices communicatively coupled via a local area network). Additionally or alternatively, such generated media modules may be stored and manipulated on one or more remote computing devices, with the resulting output transmitted to one or more remote devices for playback over a wide area network.

[0085] As shown, the process begins at block 302 and proceeds to a clock / metronome at block 304, where inputs of a tempo 306 and a time signature 308 are received. The tempo 306 and time signature 308 can be selected by the artist or can be automatically determined or generated using a model. The process proceeds to block 310, where a chord change can be triggered, receiving as input a chord change frequency parameter 312. The artist may choose to have a higher chord change frequency in music intended for a higher energy experience (e.g., dance music, uplifting atmospheric music, etc.). Conversely, a lower chord change frequency may be associated with a lower energy output (e.g., quieter music).

[0086] In block 314, a chord is selected from an available chord segment 316. A number of chord information parameters 318, 320, 322 may also be provided as inputs to the chord segment 316. These inputs may be used to determine the particular chord to be played next and output as block 324. In some instances, the artist may provide information for each chord, such as a weighting, how often that particular chord should be used, etc.

[0087] Next, in block 326, a chord variation is selected based at least in part on the harmonic complexity parameter that serves as an input. The harmonic complexity parameter 328 may be adjusted or selected by the artist, or may be automatically determined. In general, a higher harmonic complexity parameter may be associated with a higher energy audio output, and a lower harmonic complexity parameter may be associated with a lower energy audio output. In some cases, the harmonic complexity parameter may include inputs such as chord inversions, voicing, and harmonic density.

[0088] In block 330, the process gets the root of the chord and selects bass segments to play from available bass segments 334 in block 332. These bass segments then undergo bass processing 336, where equalization, filtering, timing, and other processing may be performed.

[0089] Returning to the chord variation of block 326, the process continues separately to block 338 to play a selected harmony from among the available harmony segments 340. This harmony segment then undergoes bass processing 342. Similar to low bass processing, the harmony segment bass processing 342 can include equalization, filtering, timing, etc.

[0090] Returning to the selected chord 324, the process continues to filter the melody notes individually at block 344 utilizing the input of melodic constraints 346. The output at block 348 is the melody notes available for playback. The melodic constraints 346 may be provided by the artist and may, for example, specify which notes to play or not, limit the melodic range, or provide other such constraints that may depend on the particular selected chord 324.

[0091] In block 350, the process determines which melody note (among available melody notes 348) to play. This determination can be made automatically based on model values, artist-provided input, randomization effects, or any other suitable input. In the illustrated example, one input is from trigger melody note block 352, which is based on a melodic density parameter 354. The artist can provide the melodic density parameter 354, which in part determines how complex and / or high-energy the audio output will be. Based on that parameter, melody notes may be triggered more or less frequently and at specific times using block 352, which is input to block 350, to determine which melody note to play. In various examples, the output of block 350 may be provided as an input to block 350 in the form of a feedback loop, such that the next melody note selected in block 350 depends at least in part on the melody note last selected in block 350. Next, in block 356, a melody segment is selected from available melody segments 358, after which the melody segment undergoes bass processing 360.

[0092] Returning to the start at block 302, the process separately proceeds to block 362 to play non-musical content. This may be, for example, nature sounds, speech sounds, or other such non-musical content. Various non-musical segments 364 may be stored and played. These non-musical content segments may also undergo bass processing at block 366.

[0093] The outputs of these various paths (e.g., selected bass segment(s), harmony segment(s), melody segment(s), and / or non-musical segment(s)) may each undergo separate bus processing before being combined at block 368 via a mixing and mastering process, where combined levels may be set, various filters may be applied, relative timing may be established, and any other appropriate processing steps may be performed before the generated audio content is output at block 370. In various examples, some of the paths may be omitted entirely. For example, the generative media module may omit the option to play non-musical content along with the generated musical content. The process 300 shown in FIG. 3 is exemplary only, and those skilled in the art will recognize that appropriate modifications can be made to the process 300 shown therein, and further, that there are numerous suitable alternative processes that may be used to generate generated media content.

[0094] 4 is an exemplary architecture for storing and retrieving generated media content. In this example, the generated media content includes a variety of individual tracks (each having multiple variations related to energy level or another parameter) that can be selected and played in various orders and groupings depending on certain input parameters.

[0095] As shown, generated media content 404 may be stored as one or more audio files associated with global generated media content metadata 402. Such metadata may include, for example, a global tempo (e.g., beats per minute), a global trigger frequency (e.g., how often to check for changes in an input parameter), and / or a global crossfade duration (e.g., the time to fade between different selected energies).

[0096] Within the generated media content 404 are multiple different tracks 406, 408, and 410. During operation, these tracks can be selected and played in various arrangements (e.g., randomized grouping with some kind of overlay, or played according to a predetermined sequence, etc.). In some examples, the generated media content 404, including tracks 406, 408, and 410, can be stored locally via one or more playback devices, while one or more remote computing devices can periodically transmit updated versions of the tracks, generated media content, and / or global generated media content metadata. In some examples, the remote computing devices can be periodically polled or queried by the playback devices, and in response to the queries or polls, the remote computing devices can provide updates to the generated media modules stored on the local playback devices.

[0097] For each track, there may be a corresponding subset of that track corresponding to a different energy level. For example, a first energy level (EL) for track 1 at 412, a second energy level for track 1 at 414, and an nth energy level for track 1 at 416. Each of these may include both metadata (e.g., metadata 418, 420, 422) and specific media files (e.g., media files 424, 426, 428) corresponding to the particular energy level. In some examples, each track may include multiple media files (e.g., media file 424) arranged in a particular manner, the arrangement and combination of which may be desired by the corresponding metadata (e.g., metadata 418). The media files may be in any suitable format, for example, playable via a playback device and / or streamable to a playback device for playback. In some examples, one or more of the media files 424, 426, 428 may be the output of the generative model shown in FIG. 3. The metadata may include, for example, tempo (if different from the global tempo), trigger frequency (if different from the global trigger frequency), sequence information (e.g., whether to play specific files sequentially, randomly, or with percentage weighting), crossfade duration (if different from the global crossfade), spatial information (e.g., to render audio content in space using multiple transducers), polynomial information (e.g., allowing multiple audio files to be played in this segment at once), and / or level (e.g., level adjustment in dB, or random within a predetermined range).

[0098] During operation, a target energy level can be determined using one or more input parameters (e.g., the number of people present in the room, the time of day, etc.). This determination can be made using the playback device and / or one or more remote computing devices. Based on this determination, specific media files corresponding to the determined energy level can be selected. The generative media module can then arrange and play those selected tracks according to the generative content model. This can include playing the selected tracks in a specific, predetermined order, playing them in a random or pseudo-random order, or any other suitable approach. The tracks can be played in an at least partially overlapping manner in some examples. It can be useful to vary the amount of overlap between tracks so that a casual listener does not hear a repeated loop of audio content, but instead perceives the generated audio as an endless stream of non-repeating audio.

[0099] While the example shown in FIG. 4 utilizes energy level as a parameter for distinguishing different generated audio content, in various examples, the particular variation or permutation of the generated audio content may vary along other dimensions (e.g., genre, time of day, associated user task, etc.).

[0100] d. Exemplary Sensor Data Sources and Other Input Parameters As previously discussed, the generated media module 214 may generate generated media based at least in part on input parameters, which may include sensor data (e.g., if received from a sensor data source 218) and / or other suitable input parameters. With respect to sensor input parameters, the sensor data source 218 may include data from any suitable sensor, wherever located relative to the generated media group and whatever values ​​measured thereby. Examples of suitable sensor data include physiological sensor data, such as data obtained from biometric sensors, wearable sensors, etc. Such data may include physiological parameters such as heart rate, respiration rate, blood pressure, brain waves, activity level, movement, body temperature, etc.

[0101] Suitable sensors include wearable sensors configured to be worn or carried by a user, such as a headset, a watch, a mobile device, a brain-machine interface (e.g., Neuralink), headphones, a microphone, or other similar devices. In some examples, the sensors may be non-wearable sensors or may be affixed to a fixed structure. The sensors may provide sensor data that may include, for example, data corresponding to brain activity, speech, position, movement, heart rate, pulse, body temperature, and / or sweating. In some examples, the sensors may correspond to multiple sensors. For example, as described elsewhere herein, the sensors may correspond to a first sensor worn by a first user, a second sensor worn by a second user, and a third sensor not worn by a user (e.g., affixed to a stationary body or structure). In such examples, the sensor data may correspond to multiple signals received from each of the first, second, and third sensors.

[0102] The sensors can be configured to acquire or generate information generally corresponding to the user's mood or emotional state. In one example, the sensor is a wearable brain-sensing headband, one of many examples of sensors described herein. Such a headband can include, for example, an electroencephalography (EEG) headband having multiple sensors thereon. In some examples, the headband can correspond to any of the Muse™ headbands (InteraXon; Toronto, Canada). The sensors can be positioned at various locations around the inner surface of the headband to correspond to, for example, different brain anatomies of the user (e.g., the frontal, parietal, temporal, and sphenoid bones). In this manner, each sensor can receive different data from the user. Each of the sensors can correspond to an individual channel that can be streamed from the headband to system device 210 and / or 250. Such sensor data can be used to detect the user's mood, for example, by classifying the frequency and intensity of various brain waves or by performing other analyses. Further details on using a brain-sensing headband for generating audio content can be found in commonly owned U.S. Patent Application No. 62 / 706,544, filed August 24, 2020, entitled MOOD DETECTION AND / OR INFLUENCE VIA AUDIO PLAYBACK DEVICES, which is incorporated herein by reference in its entirety.

[0103] In some examples, sensor data sources 218 include data obtained from networked device sensor data (e.g., Internet of Things (IoT) sensors, such as networked lights, cameras, temperature sensors, thermostats, presence detectors, microphones, etc.) Additionally or alternatively, sensor data sources 218 may include environmental sensors (e.g., measuring or displaying weather, temperature, time / day / week / month, etc.).

[0104] In some examples, the generated media module 214 can utilize inputs in the form of playback device capabilities (e.g., number and type of transducers, output power, other system architecture), device location (e.g., location relative to one or more users, relative to other playback devices). Further examples of generating and modifying generated audio as a result of user and device location are described in more detail in commonly owned U.S. patent application Ser. No. 62 / 956,771, filed January 3, 2020, and entitled "GENERATIVE MUSIC BASED ON USER LOCATION," which is incorporated herein by reference in its entirety. Additional inputs can include device states of one or more devices in a group, such as thermal state (e.g., if a particular device is in danger of overheating, the generated content can be modified to lower the temperature), battery level (e.g., bass output can be reduced in a portable playback device with a low battery level), and coupling state (e.g., whether a particular playback device is configured as part of a stereo pair, coupled with a sub, configured as part of a home theater setup, etc.). Any other suitable device characteristics or states may similarly be used as inputs for the generation of generated media content.

[0105] Another exemplary input parameter includes user presence; for example, when a new user enters a space playing generated audio, the user's presence can be detected (e.g., via a proximity sensor, beacon, etc.) and the generated audio can be modified in response. The modification can be based on the number of users (e.g., presenting ambient meditative audio for one user, relaxing music for two to four users, and party or dance music for more than four users). The modification can also be based on the identities of users present (e.g., user profiles based on user characteristics, listening history, or other such indicia).

[0106] In one example, a user may wear a biometric device that can measure various biometric parameters, such as the user's heart rate or blood pressure, and report those parameters to devices 210 and / or 250. The generated media modules 214 of these devices 210 and / or 250 may use these parameters to further adapt the generated audio, such as by increasing the tempo of the music in response to detecting a high heart rate (which may indicate that the user is engaged in high athletic activity) or by decreasing the tempo of the music in response to detecting high blood pressure (which may indicate that the user is stressed and could benefit from calming music).

[0107] In yet another example, one or more microphones of a playback device (e.g., microphone 115 of FIG. 1F) can detect the user's voice. The captured voice data can then be processed to, for example, determine the user's mood, age, or gender, identify a particular user among multiple users in a home, or any other such input parameter. Other examples are possible as well.

[0108] e. Examples of coordination between group members FIG. 5 is a functional block diagram illustrating data exchange in a system for playback of generated media content. For illustrative purposes, the system 500 shown in FIG. 5 includes interactions between a coordinator device 210 and a member device 250b. However, the interactions and processes described herein can be applied to interactions involving multiple additional coordinator devices 210 and / or member devices 250b. As shown in FIG. 5, the coordinator device 210 includes a generated media module 214a that receives inputs including input parameters 502 (e.g., sensor data, media content, model parameters of the generated media module 214a, or other such inputs) and clock and / or timing data 504. In various examples, the clock and / or timing data 504 can include synchronization signals for synchronizing playback and / or synchronizing generated media being generated by various devices in a group. In some examples, the clock and / or timing data 504 can be provided by an internal clock, processor, or other such component housed within the coordinator device 210 itself. In some examples, the clock and / or timing data 504 may be received from a remote computing device via a network interface.

[0109] Based on these inputs, the generative media module 214a may output generated media content 404a. Optionally, the output generated media content 404a may itself serve as input to the generative media module 214a in the form of a feedback loop. For example, the generative media module 214a may generate subsequent content (e.g., audio frames) using a model or algorithm that depends at least in part on previously generated content.

[0110] In the illustrated example, member device 250b similarly includes generated media module 214b, which may be substantially identical to generated media module 214a of coordinator device 210, or may differ in one or more aspects. Generated media module 214b may similarly receive input parameters 502 and clock and / or timing data 504. These inputs may be received from coordinator device 210, from other member devices, from other devices on the local network (e.g., a locally networked smart thermostat providing temperature data), and / or from one or more remote computing devices (e.g., a cloud server providing clock and / or timing data 504, or weather data, or any other such input). Based on these inputs, generated media module 214b may output generated media content 404b. This generated generated media content 404b may, optionally, be fed back to generated media module 214b as part of a feedback loop. In some examples, the generated media content 404b can include or consist of the generated media content 404a (generated via the coordinator device 210) that is sent over the network to the member device 250b. In other cases, the generated media content 404b can be generated independently and separately from the generated media content 404a generated via the coordinator device 210.

[0111] The generated media content 404a and 404b can then be played back via devices 210 and 250b themselves and / or for playback by other devices in the group. In various examples, the generated media content 404a and 404b can be configured for simultaneous and / or synchronous playback. In some cases, the generated media content 404a and 404b may be substantially identical or similar to one another, with each generated media module 214 utilizing the same or similar algorithms and the same or similar inputs. In other examples, the generated media content 404a and 404b may be different from one another, while still being configured for synchronized or simultaneous playback.

[0112] f. Exemplary Generative Media Using a Distributed Architecture As previously mentioned, generating media content can be computationally intensive and, in some cases, may be impractical to perform entirely on the local playback device alone. In some examples, a generated media module of the local playback device can request generated media content from generated media modules stored on one or more remote computing devices (e.g., cloud servers). The request can include or be based on specific input parameters (e.g., sensor data, user input, contextual information, etc.). In response to the request, the remote generated media module can stream the specific generated media content to the local device for playback. The specific generated media content provided to the local playback device can change over time depending on the specific input parameters, the configuration of the generated media module, or other such parameters. Additionally or alternatively, the playback device can store individual tracks for playback (e.g., different variations in the track are associated with different energy levels, as shown in FIG. 4). The remote computing device can then periodically provide new files for the updated tracks to the local playback device for playback or provide updates to the generated media module that determines when and how to play specific files stored locally on the playback device.

[0113] In this manner, the tasks required to generate and play back generated audio are distributed between one or more remote computing device(s) and one or more local playback devices. By performing at least some of the computationally intensive tasks associated with generating new media content on the remote computing devices and, optionally, by reducing the need for real-time computation, overall efficiency can be improved. By generating a discrete number of alternative tracks or track variations according to a particular media content model prior to playback via the remote computing devices, the local playback device can request and receive particular variations based on real-time or near-real-time input parameters (e.g., sensor data). For example, the remote computing devices can generate different versions of media content, and the playback device can request particular versions in real time based on the input parameters. This results in playback of appropriate generated media content based on real-time or near-real-time input parameters (e.g., sensor data) without requiring de novo generation of such media content performed in real time.

[0114] 6 is a schematic diagram of an exemplary distributed generative media playback system 600. As shown, an artist 602 can provide multiple media segments 604 and one or more generative content models 606 to a stored generative media module 214 via one or more remote computing devices. A media segment can correspond, for example, to a particular audio segment or seed (e.g., individual notes or chords, a short n-bar track, non-musical content, etc.). In some examples, the generative content model 606 can also be provided by the artist 602. This can include providing the entire model, or the artist 602 can provide input to the model 606 by, for example, modifying or adjusting certain aspects (e.g., tempo, melodic constraints, harmonic complexity parameters, chord change density parameters, etc.).

[0115] The generative media module 214 may receive both a media segment 604 and one or more input parameters 502 (as described elsewhere herein). Based on these inputs, the generative media module 214 may output generated media. As shown in FIG. 6, the artist 602 may optionally audition the generative media module 214, for example, by receiving exemplary outputs based on inputs (e.g., media segments 604 and / or generative content models 606) provided by the artist 602. In some cases, the audition may play variations of the generated media content to the artist 602 in response to a variety of different input parameters (e.g., one version corresponding to a high energy level intended to create an exciting or uplifting effect, another version corresponding to a low energy level intended to create a calming effect, etc.). Based on the output from this auditioning step, the artist 602 may dynamically update settings of the media segment 604 and / or generative content model 606 until a desired output is achieved.

[0116] In the illustrated example, there may be an iteration in block 608 every n hours (or minutes, days, etc.), during which the generated media module 214 may generate multiple different versions of the generated media content. In the illustrated example, there are three versions: version A in block 610, version B in block 612, and version C in block 614. These outputs are stored (e.g., via a remote computing device) as generated media content 616. A particular version (in this example, version C as block 618) may be sent (e.g., streamed) to the local playback device 250 for playback. In some examples, the particular version may correspond to tracks 406, 408, and 410 shown in FIG. 4.

[0117] Although three versions are shown here by way of example, in practice there may be many more versions of generated media content generated via a remote computing device. The versions may vary along several different dimensions, such as being suitable for different energy levels, for different intended tasks or activities (e.g., studying versus dancing), for different times of day, or any other suitable variation.

[0118] In the illustrated example, playback device 250 can periodically request a particular version of the generated media content from a remote computing device. Such a request can be based, for example, on user input (e.g., user selection via a controller device), sensor data (e.g., the number of people present in a room, background noise level, etc.), or other suitable input parameters. As illustrated, input parameters 502 can optionally be provided to (or detected by) playback device 250. Additionally or alternatively, input parameters 502 can be provided to (or detected by) remote computing device 106. In some examples, playback device 250 transmits the input parameters to remote computing device 106, which provides the appropriate version to playback device 250 without playback device 250 specifically requesting a particular version. g. Examples of how digital content is generated based on blockchain data

[0119] As previously mentioned, systems for generating and playing back generative media content may be able to interact with blockchain data (or data stored via other distributed ledger technologies). For example, as shown in FIG. 6 , a blockchain layer 620 may be utilized to provide data as input to other components of the generative media playback system 600, such as the playback device 250, input parameters 502, media segments 604, generative content models 606, and / or generative media modules 214. In various embodiments, the blockchain layer 620 may store data that can be used as one or more input parameters 502, data that can be included or used to obtain or affect a particular media segment 604, data that can be included or used to obtain or affect a particular generative content model 606, and / or data that can be included or used to obtain or affect a particular generative media module 214. Furthermore, some or all of these components may communicate with the blockchain layer 620 to write data to the blockchain, record transactions, or otherwise interact with the blockchain layer 620. For example, the playback device 250 may record a transaction reflecting the playback of a particular track on the blockchain layer 620. Data stored via the blockchain layer 620 may include particular input parameters 502, or data stored via the blockchain layer 620 may be used to generate appropriate input parameters 502. Similarly, particular media segments 604, generative content models 606, generative media modules 214, input parameters 502, or other appropriate data may be written to the blockchain layer 620 to create an immutable record of such content, transactions, or other data. Additional details regarding the utilization of blockchain technology (or other suitable distributed ledger technologies) in the creation and playback of generative media content are described in more detail below.

[0120] 7 is a schematic diagram of another example distributed generative media playback system 700. As shown, system 700 includes or communicates with a blockchain layer 620 to obtain, generate, or store, for example, input parameters 502, generative content models 606, or other data or parameters used in the creation and playback of generative media content.

[0121] Examples of such a blockchain layer 620 include public distributed ledgers such as Ethereum, Bitcoin, Solana, Avalanche, and Polygon. While a blockchain layer 620 is illustrated, various embodiments may use any suitable distributed ledger technology, including private or semi-private blockchains, as well as non-blockchain implementations such as directed acyclic graphs (DAGs) (e.g., Nano, IOTA, etc.). In various examples, participants using the blockchain layer 620 may transact with each other in a peer-to-peer manner, and the operation of the blockchain layer 620 may be decentralized so that no single central entity controls the operation of the network. Such distributed ledgers may be used to track the creation, exchange, and redemption of certain real-world assets, such as currencies. This approach allows for robust auditing of asset transactions due to the practical immutability of data stored on the blockchain. Currencies are just one of various assets that may be desirable to track on a distributed ledger. Other types of assets may differ from currencies with respect to one or more operations governing the creation, exchange, and / or redemption of the asset. Additionally, different blockchain architectures may differ in terms of policies, protocols, and even the tools used to program asset behavior.

[0122] Typically, a distributed ledger, in which each unit of an asset is represented by some form of digital token, can be programmed to impart a set of behaviors appropriate to the asset it represents. For example, "fungible" behavior allows an asset to be exchanged for other assets of the same class. A currency of a given denomination (e.g., $1) has the same value as other currencies of the same denomination, so all units of it are fungible. In contrast, property titles are "non-fungible" because their value depends on the size, location, and other aspects of the designated property. For each asset represented as a token, the appropriate fungible or non-fungible behavior is programmed into the class of token in the virtual ledger that tracks the asset.

[0123] In some embodiments, tokens traded via the blockchain layer 620 are non-fungible. Such non-fungible tokens (NFTs) are unique and cannot be exchanged for other tokens. NFTs may consist of and / or be associated with unique digital artwork and / or music, domain names, digital collectibles (e.g., CryptoKitties, memes, etc.), event tickets, parts of virtual worlds, digital objects used in games, avatars or characters, items with utility (e.g., providing voting or governance rights, etc.), etc. In various examples, the NFT itself may contain associated data (e.g., raw audio data for a music NFT may be stored on-chain), or the NFT may contain a pointer (e.g., a URL or URI) that directs to data stored elsewhere (e.g., audio data stored on a server controlled by the issuer of the music NFT).

[0124] Such tokens, whether fungible or non-fungible, can be stored by users via digital wallets. A digital wallet is a device, physical medium, program, or service that can store public and / or private keys for blockchain transactions. In some examples, a digital wallet can store multiple public / private key pairs for various different blockchains, allowing users to store assets related to different blockchains in a single wallet. Examples include MetaMask, Phantom, Coinbase Wallet, and Ledger Nano. In operation, users can sign blockchain transactions via their wallets using the appropriate private key (or authorize the wallet to sign transactions with the private key). If the transaction signature is valid, the transaction is then confirmed and added to the corresponding block on the blockchain. In some examples, the wallet identifier itself may be used as an input parameter to a generative model, independent of the tokens held in a particular wallet.

[0125] In various implementations, the blockchain layer 620 can be configured to automatically execute transactions under one or more conditions. Such self-executing transactions may be referred to as “smart contracts.” A smart contract is computer code stored on the blockchain and configured to execute only in certain circumstances or in certain ways. For example, a smart contract may be configured to execute a particular transaction at a certain time, or when a certain threshold is exceeded, based on one or more other transactions or other suitable criteria. In some examples, the generative media module 214 and / or the generative content model 606 may be implemented in the form of a smart contract, such that interacting with the smart contract via the blockchain layer 620 enables the smart contract to output generative media content, a generative content model, or data or instructions that can be used to generate such generative media content or a generative content model.

[0126] One organizational structure unique to blockchain is the decentralized autonomous organization (DAO). DAOs are typically community-driven entities with no central authority. Such DAOs are fully autonomous and transparent, with smart contracts providing the underlying rules and enforcing agreed-upon decisions. Community voting can be conducted by token holders using on-chain transactions. Based on the results of a particular vote, the smart contract can execute specific transactions or other code to implement the DAO members' decisions. DAOs typically issue tokens to users in exchange for currency investments, donations, or for free (e.g., through "airdrops"). Token holders typically hold a certain amount of voting power, which may be proportional to the number of tokens they hold. In some cases, token holders also receive monetary revenue, such as a share of transaction fees collected by the DAO.

[0127] In the exemplary system 700 shown in FIG. 7 , the generative media module 214 can receive a number of different inputs and, in response, output one or more generated content versions 610. These content versions 610 are stored in a generated media content store 616, from which a particular selected generated content version 618 can be selected and played back via a playback device 250 or other output device (e.g., a light component of a generated media content version 618 can cause a lighting device 702 to output light, with a particular hue, color temperature, brightness, on / off, or other pattern according to the generated media content version 618). Inputs to the generative media module 214 include a generated content model 606 and input parameters 502, similar to the approach described above with respect to FIG. 6 . As previously described, in some implementations, the generated content model 606 can be stored via the blockchain layer 620 or can be obtained from data stored via the blockchain layer 620. Similarly, one or more of the input parameters 502 can include or be based on data stored via the blockchain layer 620. For example, the blockchain data can be used by the generative media module 214 to generate an appropriate output, such as “sonification” of a data stream (e.g., a real-time feed such as a cryptocurrency price value can be made into a corresponding sound output). In some examples, the blockchain data can include data provided by one or more “oracles,” which are typical third-party services that provide external information to smart contracts (e.g., price feeds, weather data, election results, etc.). In some examples, the generative media module 606 and / or the generated media content 616 can be stored locally via the playback device 250, in which case the input parameters 502 are streamed to the playback device 250 and used to generate a new version 610 of the generated media content 616.

[0128] Additionally or alternatively, the generative media module 214 may receive as input the output of one or more smart contracts 706. In some cases, the generative media module 214 may itself take the form of a smart contract, where program code is stored on the blockchain and automatically executes under certain conditions (e.g., a user 708 interacts with a smart contract 706, and in response, particular generated media content is sent to a specified destination). In some examples, the generative media module 214 may execute locally (rather than as a smart contract on the blockchain) or via a remote server, but may communicate with the smart contract 706 to receive input parameters 502 from or provide appropriate output to the smart contract 706. For example, particular generated media content output by the generative media module 214 may be used to generate one or more NFTs via the smart contract 706. In some implementations, each particular version of generated content generated by the generative media module 214 can take on a corresponding NFT, thereby making each NFT generated by the smart contract 706 unique based on input from the generative media module 214. This is illustrated in FIG. 7, where multiple discrete NFTs 710a-f, denoted NFT1-NFTn, are generated via the smart contract 706. These NFTs 710 can then be provided as inputs to the generative media module 214. For example, the generative media module 214 can dynamically generate different content based, at least in part, on the particular NFT 710 with which it interacts. Additionally or alternatively, a user 708 can access a particular generative media module 214 only if the user holds the appropriate NFT 710 in their digital wallet.

[0129] In some examples, one or more NFTs 710 are pre-existing NFTs owned by a third-party person or entity (e.g., a person or entity not associated with user 708) and are temporarily accessible by generative media module 214. In particular examples, one or more NFTs 710 do not consist of audio data, but instead consist of another data type (e.g., video, images, or other data), which is "sonified" or otherwise converted by generative media module 214 (or another suitable component) into a form from which media content can be generated.

[0130] In the illustrated example, the smart contract 706 can also output an NFT 712, represented by NFT0, held by the user 708. Additionally, this NFT 712 can interact with or be generated through an artist DAO 714, which can communicate with one or more smart contracts 706. As previously discussed, a DAO is typically a community-driven organization in which members hold tokens (e.g., NFTs 712) that designate membership, provide voting and other governance rights, and additionally grant token holders economic benefits, such as income from future DAO revenues. In some examples, the artist DAO 714 can distribute a portion of incoming royalties (music royalties) to holders of appropriate NFTs or other tokens (e.g., the user 708 can receive economic benefits from the artist DAO 714 based, at least in part, on the user's ownership of the NFT 712).

[0131] Optionally, data corresponding to NFT 712 may be stored or embedded via a physical medium. For example, as shown in FIG. 7, data corresponding to NFT 712 may be embedded in vinyl record 716 (e.g., via a unique QR code, a code embedded in the grooves of vinyl record 716), or the data may be stored using any other suitable technology via another physical medium (e.g., an NFC or other RF tag). While vinyl record 716 is illustrated, as other examples, in various embodiments, the physical substrate may take various forms, such as a playback device, a physical card or ticket, a poster, etc.

[0132] The use of the blockchain layer 620, smart contracts 706, DAO 714, and / or NFTs 710 and 712 can provide several advantages to the generative media playback system 700. For example, by associating a specific NFT with generative media content (e.g., a soundscape), a user 708 can obtain a personalized history, which can then be transferred via the decentralized peer-to-peer transaction mechanism of the blockchain layer 620. However, a problem with this approach is that the data contained in NFTs is generally static, as opposed to the dynamic data of generative soundscapes and other generative media content. Another issue associated with NFTs is link breakage, where a locator within an NFT no longer references the artwork associated with the NFT, causing the data in the NFT to become outdated. One way to mitigate the possibility of link breakage is to store the generative media content or generative media engine on the blockchain layer 620. Furthermore, the artwork itself may be embedded in the NFT (e.g., artwork data is stored on-chain rather than on a separate server).

[0133] In some cases, an NFT 710 may include a specific seed for use by the generative media module 214. Examples of such seeds include the media segment 604 of FIG. 6, the track, energy level, or metadata of FIG. 4, or any of the various components shown in FIG. 3. Optionally, the characteristics of a particular NFT (at least with respect to use by the generative media module 14) may depend on its transaction history. For example, the specific seed associated with an NFT may change dynamically depending on when and how many times the NFT was last traded. Additionally or alternatively, different combinations of NFTs 710 connected to the generative media module 214 may result in different generated content versions 610, with the specific generative media output depending on which of the NFTs 710 are used as inputs.

[0134] In some examples, additional data related to the generative media playback system 700 can be stored via the blockchain layer 620. For example, a user's 708 listening history can be saved on the blockchain layer 620 to provide an immutable record of that listening history. This can include viewing history of generative media content or non-generative content (such as standard pre-recorded audio tracks or other content). In some cases, “followers” ​​of a particular user 708 can subscribe to that user's listening history by accessing the data stored via the blockchain layer 620. Because blockchains are generally permissionless and transparent, followers have free access to the listening history (or other content data associated with a particular network address). In yet another example, followers can subscribe to a particular user's 708 generated media content. This allows the user's 708 generated media content to be dynamically created based on various inputs, but this same media content can be enjoyed by other followers.

[0135] For example, consider an artist who wants to create a particular soundscape using a generative media module 214. Fans of the artist can listen to the soundscape in real time or near real time via data on the blockchain. In some cases, followers utilize their own local generated media modules 214 (or those running on other devices) to generate corresponding generated media content using inputs, pointers, or other data. In at least some examples, such local generated media modules 214 may also utilize additional local inputs (e.g., specific playback device characteristics, local sensor data, etc.). As a result, the artist's generated media content is merged with that generated by the user's own local generated media module 214, resulting in generated media content that reflects the artist's intent but is somewhat modified. Optionally, followers of the artist may be required to hold specific NFTs or other tokens to access the artist's generated media content. In yet another example, a particular playlist or radio station may be accessible only to users who hold a particular NFT or token.

[0136] h. Exemplary Methods for Generating and Playing Generative Audio 8-13 are flow diagrams of example methods for playing generated audio content via multiple separate playback devices. Methods 800, 900, 1000, 1100, 1200, and 1300 may be implemented by any of the devices or systems described herein, or any other device or system now known or later developed.

[0137] The various examples of methods 800, 900, 1000, 1100, 1200, and 1300 include one or more operations, functions, or actions illustrated by blocks. While the blocks are shown in sequential order, these blocks may also be performed in parallel and / or in orders different from those disclosed and described herein. Additionally, various blocks may be combined into fewer blocks, divided into additional blocks, and / or eliminated based on the desired implementation.

[0138] Furthermore, for methods 800, 900, 1000, 1100, 1200, and 1300, as well as other processes and methods disclosed herein, flowcharts illustrate the functionality and operation of some example possible implementations. In this regard, each block may represent a module, segment, or portion of program code, including one or more instructions executable by one or more processors to implement specific logical functions or steps in the process. The program code may be stored on any type of computer-readable medium, such as a storage device including a disk or hard drive. Computer-readable media may include non-transitory computer-readable media, such as tangible non-transitory computer-readable media for storing short-term data, such as register memory, processor cache, and random access memory (RAM). Computer-readable media may also include non-transitory media, such as secondary or persistent long-term storage devices, such as read-only memory (ROM), optical or magnetic disks, and compact disc read-only memory (CD-ROM). Computer-readable media may also be any other volatile or non-volatile storage system. The computer-readable medium may be considered, for example, a computer-readable storage medium or a tangible storage device. Furthermore, for the methods and other processes and methods disclosed herein, each block in Figures 8-13 may correspond to circuitry hardwired to perform specific logical functions within the process.

[0139] 8, method 800 begins at block 802 with receiving a command to play generated media content via a group of playback devices or combined zones. Such a command may be received, for example, via control device 130 or other suitable user input.

[0140] At block 804, the method 800 includes the group coordinator device providing timing information to the producing group member devices. The timing information may include contextual timing data (e.g., time data associated with sensor inputs or other user inputs), produced media playback timing data (e.g., timestamps and synchronization data to facilitate synchronized playback of the produced media), and / or media content stream timing data based on a common clock.

[0141] At block 806, the method optionally includes determining a generated media content model to be used to generate the generated media. Such a model may be implemented, for example, in media content module 214 described above with respect to FIGS. 2-6. In some examples, each of the member devices may utilize the same or substantially the same generated media content model, while in other cases, some or all of the member devices may utilize different generated media content models. For example, a first generated media content model may generate rhythmic beats, while a second generated media content model may generate ambient nature sounds. When played simultaneously, the generated audio produced by these different generated media content models may create a pleasant listening experience for the user. In some examples, the selection of a particular generated media content model may itself be based on one or more input parameters, such as device capabilities, device location, number of users present, user sensor data, etc.

[0142] At block 808, method 800 includes the coordinator device and the member devices receiving context and / or other input data. For example, the input data may include sensor data, user input, context data, or any other relevant data that may be utilized as input for the generated media content model.

[0143] The method 800 continues at block 810 with the coordinator device and the member devices synchronizing to generate and play the generated media content.

[0144] 9 illustrates another method 900 for playing generated audio content through multiple playback devices. Method 900 begins with receiving one or more input parameters at a group coordinator device at block 902. As previously mentioned, the input parameters may include sensor data, user input, contextual data, or any other input that may be used by the generated media module to generate generated audio for playback.

[0145] In block 904, the coordinator device transmits input parameters to one or more discrete playback devices having the generated media module. For example, the coordinator device may obtain sensor data and other input parameters and transmit them to multiple discrete playback devices in the environment or to multiple discrete playback devices distributed across multiple environments. In some examples, these input parameters may include features of the generated content model itself, for example, providing instructions for updating the generated media module stored locally by one or more of the discrete playback devices.

[0146] At block 906, the method includes transmitting timing data from the coordinator device to the playback devices. The timing data may include, for example, clock data or other synchronization signals configured to facilitate coordination of the generation of the generated media content and synchronized playback of that generated media content via the separate playback devices.

[0147] The method 900 continues at block 908 with simultaneously playing the generated media content via the playback devices based at least in part on the input parameters. As previously mentioned, the various playback devices may play the same generated audio, or each may play separate generated audio that, when played in sync, produces a desired psychoacoustic effect for the user present.

[0148] In the example of Figure 9, the generated media content may be generated locally by discrete playback devices, each generating and playing their own generated audio content in parallel with each other. In another method 1000 shown in Figure 10, the generated media content is generated at a coordinator device, which then sends the generated media content along with timing data to separate playback devices for synchronized playback.

[0149] At block 1002, the method 1000 includes receiving, at the group coordinator device, one or more input parameters. Examples of the input parameters are described elsewhere herein and include sensor data, user input, contextual data, or any other input that can be used by the generated media module to generate generated audio for playback.

[0150] At block 1004, the coordinator device generates first and second generated media streams based at least in part on the input parameters, and at block 1006, the first and second media streams are transmitted to first and second separate playback devices, respectively. For example, the coordinator device may generate two streams forming different channels of generated audio, e.g., with a left channel played by the first playback device and a corresponding right channel played by the second playback device. Additionally or alternatively, the two streams may be separate audio tracks that can nevertheless be played synchronously, such as a rhythmic beat in one stream and ambient nature sounds in the other stream. Multiple other variations are possible. While this example describes two streams for two playback devices, in various other examples, there may be one stream or more than two streams that can be provided to any number of playback devices for synchronous playback. In at least some examples, one or more of the playback devices may be located in different environments (e.g., different homes, different cities, etc.) that are far removed from one another.

[0151] In block 1008, the first playback device plays the first originating media stream and the second playback device now plays the second originating media stream. In some examples, this simultaneous playback can be facilitated by using timing data received from the coordinator device.

[0152] 11 shows another exemplary method 1100 for generating and playing back generated media content. As previously mentioned, it may be beneficial to use one or more remote computing devices (e.g., cloud-based servers) to perform at least a portion of the processing necessary to generate the generated media content to reduce the computational demands placed on the local playback device and / or to perform operations that are not feasible using components of the local playback device. Method 1100 begins at block 1102 with receiving one or more input parameters at the playback device. As previously mentioned, the input parameters may include sensor data, user input, contextual data, or any other input that may be used by the generated media module to generate generated audio for playback.

[0153] At block 1104, method 1100 includes accessing a library containing a plurality of existing media segments. For example, a plurality of individual media segments (e.g., audio tracks) may be stored on a playback device and arranged and / or mixed for playback according to a generative content model. Additionally or alternatively, the library may be stored on one or more remote computing devices, and the individual media segments may be transmitted from the remote computing devices to the playback device for playback.

[0154] The method 1100 continues at block 1106 with generating media content based at least in part on the input parameters by arranging a selection of existing media segments from a library for playback according to the generated media content model. As described elsewhere herein, the generated media content model may receive one or more input parameters as input. Based on the input, the generated media content model may be used to output particular generated media content. In examples, the generated media content may include an arrangement of existing media segments, for example, arranging them in a particular order, with or without overlap between the particular media segments, and / or with additional processing or mixing steps performed to generate the desired output.

[0155] At block 1108, the playback device plays the generated media content. In various examples, this playback can be performed simultaneously and / or synchronously with additional playback devices.

[0156] FIG. 12 illustrates another example method 1200 for generating and playing back generative media content. As discussed above, incorporating or relying on blockchain data to create generative media content can be beneficial. Method 1200 begins, at block 1202, with accessing, via a playback device, blockchain data stored via a distributed ledger. The distributed ledger can be a public blockchain, such as Ethereum, Bitcoin, or Solana, or optionally a private or semi-private blockchain or non-blockchain ledger. The blockchain data can include one or more pre-existing media segments or other seeds used to generate the generative media content. In some examples, such data is stored directly on the blockchain itself, while in other examples, the blockchain can store pointers (e.g., URLs or URIs) that indicate the stored locations of the media segments or other seed data. Optionally, the blockchain data takes the form of one or more non-fungible tokens (NFTs).

[0157] At block 1204, method 1200 generates media content via a playback device based at least in part on data on the blockchain. In some examples, this generation includes accessing a library of existing media segments stored on the playback device or other suitable storage location (e.g., a remote server, another device on a local network, etc.). In some cases, these media segments are retrieved from the blockchain or other remote location and saved via the playback device. The playback device can then arrange for playback of a selection of existing media segments from the library in accordance with the generative media content model. This selection can be based at least in part on data on the blockchain. For example, as described elsewhere herein, the generative media content model can communicate with a smart contract or decentralized autonomous organization (DAO) in a manner that influences the particular generative media content output by the model. In some cases, particular NFTs or other tokens can influence the output of the generative media content model. Such NFTs or other tokens can, in some cases, be used in combination, such that a particular combination of NFTs or other tokens generates a unique output via the generative media content model. In at least some instances, two or more blockchains may be utilized simultaneously in this manner (e.g., a generative media content model may vary output based on a user holding both a first NFT on the Solana network and a second NFT on the Ethereum network). At block 1206, method 1200 includes playing the generated media content via a playback device.

[0158] FIG. 13 illustrates another exemplary method 1300 for generating and playing generative media content. As described above, smart contracts or other self-executing code can be used to generate, store, and play generative media content. The method 1300 begins in block 1302 with transmitting, via a playback device, data associated with a first token over a network to a network address in a distributed ledger. The address can be associated with a generative media smart contract configured to generate a generative media content model. In some examples, the first token can be an NFT, and optionally, multiple such tokens are transmitted to the address of the smart contract.

[0159] At block 1304, method 1300 receives, via the playback device, a generated media content model from a network address associated with the generated media smart contract. For example, when the smart contract executes, it generates a particular generated media content model based at least in part on data associated with the first token. This generated media content model can be provided to a user, thereby generating newly created media content based on the first token data. In addition to the first token data, the smart contract can also generate different outputs based on other input parameters (e.g., sensor data, playback device characteristic data, playback device state, user listening history data, etc.), as described elsewhere, provided such data is provided to the smart contract address. Additionally, or alternatively, the generated media content model provided to the user can output different media content based on one or more other input parameters.

[0160] Next, at block 1306, the playback device generates media content based at least in part on the generated media content model. In some examples, this generation includes accessing a library of existing media segments stored on the playback device or other suitable storage location (e.g., a remote server, another device on a local network, etc.). The playback device can then arrange selected existing media segments from the library for playback according to the generated media content model. At block 1308, method 1300 plays the generated media content via the playback device.

[0161] Various examples of generated media playback are described herein. Those skilled in the art will appreciate that a wide variety of generated media modules, algorithms, inputs, sensor data, and playback device configurations are contemplated and may be used in accordance with the present technology.

[0162] IV. Conclusion The above discussion of playback devices, controller devices, playback zone configurations, and media content sources provides only a few examples of operating environments in which the features and methods described below may be implemented. Other operating environments and configurations of media playback systems, playback devices, and network devices not explicitly described herein may also be applicable and suitable for implementation of the features and methods.

[0163] The above description discloses various exemplary systems, methods, apparatus, and articles of manufacture that include, among other things, firmware and / or software executing on hardware. It is understood that such examples are merely illustrative and should not be considered limiting. For example, it is contemplated that any or all of the firmware, hardware, and / or software aspects or components could be embodied exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Thus, the examples provided are not the only ways to implement such systems, methods, apparatus, and / or articles of manufacture.

[0164] Furthermore, references herein to an "example" mean that a particular feature, structure, or characteristic described in connection with the example may be included in at least one example or embodiment of the present invention. Appearances of this phrase in various places throughout the specification do not necessarily all refer to the same example, nor are they mutually exclusive, separate, or alternative examples from other examples. Thus, it is explicitly and implicitly understood by those skilled in the art that examples described herein can be combined with other examples.

[0165] This specification is presented primarily in terms of example environments, systems, procedures, steps, logic blocks, processes, and other symbolic representations that directly or indirectly resemble the operations of network-coupled data processing devices. These process descriptions and representations are typically used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art. Numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, those skilled in the art will understand that certain examples of the present technology may be practiced without the specific specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail to avoid unnecessarily obscuring aspects of the examples. Accordingly, the scope of the present disclosure is defined by the appended claims, rather than by the foregoing description of the illustrative embodiments.

[0166] If any of the appended claims are read to cover purely software and / or firmware embodiments, at least one of the elements in at least one example is expressly defined hereby to include a tangible, non-transitory medium, such as a memory, DVD, CD, Blu-ray®, etc., that stores the software and / or firmware.

[0167] The disclosed technology is illustrated, for example, according to various examples described below. Various examples of embodiments of the disclosed technology are described as numbered examples (1, 2, 3, etc.) for convenience. These are provided as examples and are not intended to limit the disclosed technology. It should be noted that any of the subordinate examples may be combined in any combination or may be included in their own independent examples. Other examples may be presented as well.

[0168] Example 1: A method including receiving input parameters at a coordinator device; transmitting the input parameters from the coordinator device to a plurality of playback devices, each having a generated media module therein; and transmitting timing data from the coordinator device to the plurality of playback devices such that the playback devices simultaneously play the generated media content based at least in part on the input parameters.

[0169] Example 2: The method of any one of the examples herein, wherein the first and second playback devices each play different generated audio content based at least in part on the input parameters.

[0170] Example 3: The method of any one of the examples herein, wherein the input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device state (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data).

[0171] Example 4: The method of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0172] Example 5: The method of any one of the examples herein, further comprising sending a signal from the coordinator device to at least one of the plurality of playback devices that causes the playback device to modify a generated media module.

[0173] Example 6: The method of any one of the examples herein, wherein the generated media content includes at least one of generated audio content or generated visual content.

[0174] Example 7: The method of any one of the examples herein, wherein the generating media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0175] Example 8: A device comprising: a network interface; one or more processors; and a tangible, non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the device to perform operations, the operations including: receiving input parameters via the network interface; sending the input parameters via the network interface to a plurality of playback devices, each having a generated media module therein; and sending timing data to the plurality of playback devices via the network interface such that the playback devices simultaneously play the generated media content based at least in part on the input parameters.

[0176] Example 9: The device of any one of the examples herein, wherein the first and second playback devices each play different generated audio content based at least in part on the input parameters.

[0177] Example 10: The input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data) of any one of the devices in the examples herein.

[0178] Example 11: The device of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0179] Example 12: The device of any one of the examples herein, wherein the operations further include sending a signal from the coordinator device to at least one of the plurality of playback devices via a network interface, causing the generated media module of the playback device to be modified.

[0180] Example 13: The device of any one of the examples herein, wherein the generated media content includes at least one of generated audio content or generated visual content.

[0181] Example 14: The device of any one of the examples herein, wherein the generating media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0182] Example 15: A tangible, non-transitory computer-readable medium storing instructions that, when executed by one or more processors of the device, cause the device to perform operations including: receiving input parameters at a coordinator device; transmitting the input parameters from the coordinator device to a plurality of playback devices, each having a generated media module therein; and transmitting timing data from the coordinator device to the plurality of playback devices such that the playback devices simultaneously play the generated media content based at least in part on the input parameters.

[0183] Example 16: The computer-readable medium of any one of the examples herein, wherein the first and second playback devices each play different generated audio content based at least in part on the input parameters.

[0184] Example 17: The computer-readable medium of any one of the examples herein, wherein the input parameters include one or more of: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data).

[0185] Example 18: The computer-readable medium of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0186] Example 19: The computer-readable medium of any one of the examples herein, further comprising sending a signal from the coordinator device to at least one of the plurality of playback devices that causes the playback device to modify a produced media module.

[0187] Example 20: The computer-readable medium of any one of the examples herein, wherein the generated media content includes at least one of generated audio content or generated visual content.

[0188] Example 21: The computer-readable medium of any one of the examples herein, wherein the generative media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0189] Example 22: A method comprising: receiving input parameters at a coordinator device; generating first and second media content streams via a generate media module of the coordinator device; transmitting the first media content stream to a first playback device via the coordinator device; and transmitting the second media content stream to a second playback device via the coordinator device such that the first and second media content streams are played simultaneously via the first and second playback devices.

[0190] Example 23: The method of any one of the examples herein, further comprising transmitting timing data from the coordinator device to each of the first and second playback devices.

[0191] Example 24: The method of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0192] Example 25: The method of any one of the examples herein, wherein the first and second media content streams are different.

[0193] Example 26: The method of any one of the examples herein, wherein the input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data).

[0194] Example 27: The method of any one of the examples herein, further comprising modifying the originating media module of the coordinator device.

[0195] Example 28: The method of any one of the examples herein, wherein each of the first and second generated media content streams includes at least one of generated audio content or generated visual content.

[0196] Example 29: The method of any one of the examples herein, wherein the generating media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0197] Example 30: A device comprising: a network interface; a generated media module; one or more processors; and a tangible, non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the device to perform operations, the operations including receiving input parameters via the network interface; generating first and second media content streams via the generated media module; transmitting the first media content stream to a first playback device via the network interface; and transmitting the second media content stream to a second playback device via the network interface such that the first and second media content streams are played simultaneously via the first and second playback devices.

[0198] Example 31: The device of any one of the examples herein, wherein the operations further include transmitting the timing data to each of the first and second playback devices via the network interface.

[0199] Example 32: The device of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0200] Example 33: The device of any one of the examples herein, wherein the first and second media content streams are different.

[0201] Example 34: The input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data) for any one of the devices in the examples herein.

[0202] Example 35: The device of any one of the examples herein, wherein the operations further include modifying the generated media module.

[0203] Example 36: The device of any one of the examples herein, wherein each of the first and second generated media content streams includes at least one of generated audio content or generated visual content.

[0204] Example 37: The device of any one of the examples herein, wherein the generating media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0205] Example 38: A tangible, non-transitory computer-readable medium storing instructions that, when executed by one or more processors of the coordinator device, cause the coordinator device to perform operations, the operations including: receiving input parameters at the coordinator device; generating first and second media content streams via a generating media module of the coordinator device; transmitting the first media content stream to a first playback device via the coordinator device; and transmitting the second media content stream to a second playback device via the coordinator device such that the first and second media content streams are played simultaneously via the first and second playback devices.

[0206] Example 39: The computer-readable medium of any one of the examples herein, further comprising transmitting timing data from the coordinator device to each of the first and second playback devices.

[0207] Example 40: The computer-readable medium of any one of the examples herein, wherein the timing data includes at least one of clock data or one or more synchronization signals.

[0208] Example 41: The computer-readable medium of any one of the examples herein, wherein the first and second media content streams are different.

[0209] Example 42: The computer-readable medium of any one of the examples herein, wherein the input parameters include one or more of: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity), user mood data).

[0210] Example 43: The computer-readable medium of any one of the examples herein, wherein the operations further include modifying the generated media module of the coordinator device.

[0211] Example 44: The computer-readable medium of any one of the examples herein, wherein each of the first and second generated media content streams includes at least one of generated audio content or generated visual content.

[0212] Example 45: The computer-readable medium of any one of the examples herein, wherein the generative media module includes an algorithm that automatically generates new media output based on input including at least the input parameters.

[0213] Example 46: A playback device comprising: one or more amplifiers configured to drive one or more audio transducers; one or more processors; and a data storage device having instructions that, when executed by the one or more processors, cause the playback device to perform operations, the operations including: receiving one or more first input parameters at the playback device; generating first media content via the playback device based at least in part on the one or more first input parameters, the first input parameters including: accessing a library stored on the playback device that includes a plurality of pre-existing media segments, and arranging a first selection of the pre-existing media segments from the library for playback based at least in part on the one or more input parameters according to a generated media content model; and playing the first generated media content via the one or more amplifiers.

[0214] Example 47: The playback device of any one of the examples herein, wherein the operations include receiving one or more second input parameters at the playback device that differ from the first input parameters; generating second media content via the playback device based at least in part on the one or more second input parameters, where the second media content differs from the first media content, and the generating includes accessing a library and arranging a second selection of existing media segments from the library for playback based at least in part on the one or more second input parameters according to a generated media content model; and playing the second generated media content via one or more amplifiers.

[0215] Example 48: The playback device of claim 1, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments at least partially offset in time.

[0216] Example 49: The playback device of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments so that they at least partially overlap in time.

[0217] Example 50: The playback device of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes applying different equalization adjustments to different existing media segments.

[0218] Example 51: The playback device of any one of the examples herein, wherein the step of placing a first selection of existing media segments from a library or playing includes applying gain levels that vary over time to different existing media segments.

[0219] Example 52: The playback device of any one of the examples herein, wherein the step of arranging the first selection of an existing media segment from the library or playback includes randomizing a starting point for playback of the particular existing media segment.

[0220] Example 53: The playback device of any one of the examples herein, wherein the first generated media content and the second generated media content each include new media content.

[0221] Example 54: The playback device of any one of the examples herein, wherein the first generated media content includes audio content and the plurality of pre-existing media segments includes a plurality of pre-existing audio segments.

[0222] Example 55: The playback device of any one of the examples herein, wherein the first generated media content includes audio-visual content and the plurality of existing media segments includes a plurality of existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0223] Example 56: The playback device of any one of the examples herein, further including receiving additional existing media segments via the network interface and updating the library to include at least the additional existing media segments.

[0224] Example 57: The playback device of any one of the examples herein, wherein the first and second input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity, speech characteristics), user mood data).

[0225] Example 58: A method comprising: receiving one or more first input parameters at a playback device; generating first media content via the playback device based at least in part on the one or more first input parameters, the generating step comprising: accessing a library stored on the playback device that includes a plurality of pre-existing media segments, and arranging a first selection of the pre-existing media segments from the library for playback based at least in part on the one or more input parameters according to a generated media content model; and playing the first generated media content via the playback device.

[0226] Example 59: The method of any one of the examples herein, comprising: receiving one or more second input parameters at a playback device that differ from the first input parameters; generating second media content via the playback device based at least in part on the one or more second input parameters, wherein the second media content differs from the first media content, and the generating comprises accessing a library and arranging a second selection of existing media segments from the library for playback based at least in part on the one or more second input parameters according to a generated media content model; and playing the second generated media content via the playback device.

[0227] Example 60: The method of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments to be at least partially offset in time.

[0228] Example 61: The method of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments so that they are at least partially overlapping in time.

[0229] Example 62: The method of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes applying different equalization adjustments to different existing media segments.

[0230] Example 63: The method of any one of the examples herein, wherein the step of arranging a first selection of existing media segments from a library or playing includes applying gain levels that vary over time to different existing media segments.

[0231] Example 64: The method of any one of the examples herein, wherein the step of arranging the first selection of an existing media segment from the library or playing includes randomizing a starting point for playing the particular existing media segment.

[0232] Example 65: The method of any one of the examples herein, wherein the first generated media content and the second generated media content each include new media content.

[0233] Example 66: The method of any one of the examples herein, wherein the first generated media content includes audio content and the plurality of pre-existing media segments includes a plurality of pre-existing audio segments.

[0234] Example 67: The method of any one of the examples herein, wherein the first generated media content includes audio-visual content and the plurality of existing media segments includes a plurality of existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0235] Example 68: The method of any one of the examples herein, further including receiving additional existing media segments via the network interface and updating the library to include at least the additional existing media segments.

[0236] Example 69: The method of any one of the examples herein, wherein the first and second input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity, speech characteristics), user mood data).

[0237] Example 70: A tangible, non-transitory computer-readable medium storing instructions that, when executed by one or more processors of the playback device, cause the playback device to perform operations including: receiving one or more first input parameters at the playback device; generating first media content via the playback device based at least in part on the one or more first input parameters, the generating first media content including accessing a library stored on the playback device that includes a plurality of pre-existing media segments, and arranging a first selection of the pre-existing media segments from the library for playback based at least in part on the one or more input parameters according to a generated media content model; and playing the first generated media content via the playback device.

[0238] Example 71: The computer-readable medium of any one of the examples herein, wherein the operations include receiving one or more second input parameters at a playback device that differ from the first input parameters; generating second media content via the playback device based at least in part on the one or more second input parameters, where the second media content differs from the first media content, and the generating includes accessing a library and arranging a second selection of existing media segments from the library for playback based at least in part on the one or more second input parameters according to a generated media content model; and playing the second generated media content via one or more amplifiers.

[0239] Example 72: The computer-readable medium of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments to be at least partially offset in time.

[0240] Example 73: The computer-readable medium of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes arranging two or more of the existing media segments to at least partially overlap in time.

[0241] Example 74: The computer-readable medium of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes applying different equalization adjustments to different existing media segments.

[0242] Example 75: The computer-readable medium of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes applying gain levels that vary over time to different existing media segments.

[0243] Example 76: The computer-readable medium of any one of the examples herein, wherein arranging a first selection of existing media segments from a library for playback includes randomizing a starting point for playback of the particular existing media segment.

[0244] Example 77: The computer-readable medium of any one of the examples herein, wherein the first generated media content and the second generated media content each include new media content.

[0245] Example 78: The computer-readable medium of any one of the examples herein, wherein the first generated media content includes audio content and the plurality of pre-existing media segments includes a plurality of pre-existing audio segments.

[0246] Example 79: The computer-readable medium of any one of the examples herein, wherein the first generated media content includes audio-visual content and the plurality of existing media segments includes a plurality of existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0247] Example 80: The computer-readable medium of any one of the examples herein, further including receiving additional existing media segments via the network interface and updating the library to include at least the additional existing media segments.

[0248] Example 81: The computer-readable medium of any one of the examples herein, wherein the first and second input parameters include one or more of physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiration rate, brain waves)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is combined with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiration rate, brain activity, speech characteristics), user mood data).

[0249] Example 82: A system comprising a first playback device and a second playback device, the first playback device comprising a first network interface, one or more first processors, and a data storage device having instructions having instructions that, when executed by the one or more processors, cause the first playback device to perform operations, the operations including receiving one or more input parameters, and generating media content based at least in part on the one or more input parameters, the generated media content including a first portion and at least a second portion, the generating including accessing a library stored on the playback device containing a plurality of pre-existing media segments, and arranging a selection of the pre-existing media segments from the library for playback based at least in part on the one or more input parameters according to a generated media content model; transmitting a signal including a second portion of the generated media content and corresponding timing information via an interface; and causing playback of a first portion of the generated media content, wherein the second playback device comprises a second network interface, one or more audio transducers, one or more second processors, and a data storage device having instructions having instructions that, when executed by the one or more second processors, cause the second playback device to perform operations, the operations including receiving the signal transmitted from the first playback device via the second network interface; and playing, via the one or more transducers, the second portion of the generated media content in accordance with the timing information substantially synchronized with the playback of the first portion of the generated media content.

[0250] Example 83: The system of any one of the examples herein, further comprising a network device, the network device comprising a third network interface, one or more processors, and a data storage device having instructions having instructions that, when executed by the one or more processors, cause the third playback device to perform operations, the operations including: receiving a request from the first playback device via the third network interface over a data network; and, in response to receiving the request, transmitting an updated library of existing media segments to the first playback device via the third network interface over the data network.

[0251] Example 84: The system of any one of the examples herein, wherein the network device comprises one or more of a remote server, another playback device, a mobile computing device, a laptop, or a tablet.

[0252] Example 85: A system comprising a first playback device and a second playback device communicatively coupled via a local area network, the first playback device comprising one or more first processors, one or more first audio transducers, and a data storage device having instructions having instructions that, when executed by the one or more first processors, cause the first playback device to perform operations, the operations including: receiving one or more input parameters; generating first media content based at least in part on the one or more input parameters, the first playback device accessing a first library stored on the first playback device comprising a plurality of pre-existing media segments, and arranging a selection of the pre-existing media segments from the first library for playback based at least in part on the one or more input parameters according to a first generated media content model; and playing the first generated media content via the one or more first audio transducers; and the second playback device comprising a second network interface.

[0253] 1. A system comprising: one or more second audio transducers; one or more second processors; and a data storage device having instructions having instructions, when executed by the one or more second processors, that cause a second playback device to perform operations, the operations comprising: generating second media content based at least in part on one or more input parameters, the second generated media content being substantially identical to the first generated media content, the generating comprising: accessing a second library stored on the second playback device comprising a plurality of pre-existing media segments, and arranging a selection of the pre-existing media segments from the second library for playback based at least in part on the one or more input parameters according to a second generated media content model; and playing the second generated media content via the one or more second audio transducers in synchronization with playback of the first generated media content via the first playback device.

[0254] Example 86: The system of any one of the examples herein, wherein the first produced media content model and the second produced media content model are substantially identical.

[0255] Example 87: The system of any one of the examples herein, wherein the first library and the second library are substantially identical.

[0256] Example 88: A method comprising the steps of: accessing, via a playback device, blockchain data stored on a distributed ledger; and generating, via the playback device, media content based at least in part on the blockchain data; wherein the generating step comprises: accessing a library stored on the playback device, the library including a plurality of existing media segments, and arranging a selection of the existing media segments from the library for playback in accordance with a generated media content model and based at least in part on the blockchain data; and further playing, via the playback device, the generated media content.

[0257] Example 89: The method of any one of the examples herein, wherein the NFT data includes one or more pre-existing media segments and accessing the NFT data includes storing the one or more pre-existing media segments in a library.

[0258] Example 90: The method of any one of the examples herein, wherein the data on the blockchain includes first non-fungible token (NFT) data, wherein the distributed ledger is a first distributed ledger, and the method further includes accessing, via the playback device, data related to a second NFT stored in a second distributed ledger, and wherein placing a selection of existing media segments from the library for playback according to the generative media content model is based at least in part on both the first NFT data and the second NFT data.

[0259] Example 91: Any one of the methods herein, wherein the first distributed ledger is associated with a first blockchain layer and the second distributed ledger is associated with a second blockchain layer that is different from the first blockchain layer.

[0260] Example 92: The method of any one of the examples herein, wherein the data in the blockchain is associated with a playlist.

[0261] Example 93: The method of any one of the examples herein, wherein the data on the blockchain depends, at least in part, on transactions recorded on a distributed ledger that includes non-fungible tokens (NFTs).

[0262] Example 94: The method of any one of the examples herein, wherein arranging the selection of existing media segments from the library according to the generative media content model is further based at least in part on one or more input parameters.

[0263] Example 95: The method of any one of the examples herein, wherein the input parameters include one or more of physiological sensor data, networked device sensor data, environmental data, playback device characteristic data, playback device state, user listening history data, oracle data stored via a distributed ledger, or user data.

[0264] Example 96: The method of any one of the examples herein, wherein user listening history data is stored via a distributed ledger.

[0265] Example 97: The method of any one of the examples herein, wherein accessing data on the blockchain includes connecting to a user wallet that holds a non-fungible token (NFT).

[0266] Example 98: The method of any one of the examples herein, wherein accessing data on the blockchain includes accessing a code associated with a physical media object (e.g., a QR code or other code imprinted on custom vinyl or other media) via a control device.

[0267] Example 99: The method of any one of the examples herein, wherein arranging a selection of existing media segments from a library for playback includes arranging two or more of the existing media segments in an at least partially temporally offset manner.

[0268] Example 100: The method of any one of the examples herein, wherein arranging the selection of existing media segments from the library for playback includes arranging two or more of the existing media segments to at least partially overlap in time.

[0269] Example 101: A method comprising: sending, via a playback device, over a network, data associated with a first token to a network address of a distributed ledger, the address being associated with a generative media smart contract configured to generate a generative media content model; receiving, via the playback device, the generative media content model from the network address associated with the generative media smart contract; and generating, via the playback device, media content based at least in part on the generative media content model; wherein the generating comprises: accessing a library including a plurality of existing media segments; and arranging a selection of the existing media segments from the library for playback in accordance with the generative media content model, the method further comprising playing, via the playback device, the generated media content.

[0270] Example 102: The method of any one of the examples herein, wherein the token data includes first non-fungible token (NFT) data, the method further comprising: sending, via a playback device, data relating to a second NFT stored in a distributed ledger to a network address associated with a generated media smart contract; receiving, via a playback device, a second generated media content model from the network address associated with the generated media smart contract, wherein the second generated media content model differs from the first generated media content model; generating, via the playback device, second media content based at least in part on the second generated media content model; and playing, via the playback device, the second generated media content.

[0271] Example 103: The method of any one of the examples herein, wherein token data is associated with a curated playlist.

[0272] Example 104: The method of any one of the examples herein, wherein the token data depends at least in part on transactions recorded on a distributed ledger that includes the token.

[0273] Example 105: The method of any one of the examples herein, wherein arranging the selection of existing media segments from the library according to the generated media content model is further based at least in part on one or more input parameters.

[0274] Example 106: The method of any one of the examples herein, wherein the input parameters include one or more of physiological sensor data, networked device sensor data, environmental data, playback device characteristic data, playback device state, user listening history data, oracle data stored via a distributed ledger, or user data.

[0275] Example 107: The method of any one of the examples herein, wherein the user's listening history data is stored via a distributed ledger.

[0276] Example 108: The method of any one of the examples herein, further including, prior to transmitting the token data, connecting to a user wallet that stores the token data to access the token data.

[0277] Example 109: The method of any one of the examples herein, further including, prior to transmitting the token data, accessing the token data via a code associated with the physical media object via the control device (e.g., a QR code or other code imprinted on custom vinyl or other media).

[0278] Example 110: The method of any one of the examples herein, wherein the step of arranging the selection of existing media segments from the library for playback includes arranging two or more of the existing media segments in an at least partially temporally offset manner.

[0279] Example 111: The method of any one of the examples herein, wherein the step of arranging the selection of existing media segments from the library for playback includes arranging two or more of the existing media segments to at least partially overlap in time.

[0280] Example 112: The method of any one of the examples herein, wherein the first generated media content and the second generated media content each include new media content.

[0281] Example 113: One or more tangible, non-transitory computer-readable media storing instructions for execution by one or more processors that cause a media playback system or playback device to perform an operation according to any one of the method examples herein.

[0282] Example 114: A media playback system comprising one or more processors; and a computer-readable medium bearing the method of any one of the examples herein.

[0283] Example 115: A playback device comprising one or more processors; and a computer-readable medium carrying the method of any one of the examples herein.

Claims

1. 1. A computing system comprising: a network interface; one or more processors; a memory that stores instructions for execution by one or more processors; The instructions cause the computing system to perform operations including: receiving first data having one or more input parameters via a network interface; generating synthetic content via one or more generative machine learning models based on second data and the received first data, wherein the second data comprises data corresponding to data obtained from a blockchain-based distributed ledger, and wherein the one or more generative machine learning models comprise a first generative machine learning model configured to output a first type of media content and a second generative machine learning model configured to output a second type of media content, and further wherein the generated synthetic content comprises a mixture of outputs of the first and second generative machine learning models; Transmitting the generated composite content to a network device via a network interface.

2. 2. The computing system of claim 1, wherein the first data is audio data.

3. 10. The computing system of claim 1, wherein the first data corresponds to sensor data received via one or more sensors.

4. 4. The computing system of claim 3, wherein the sensor data comprises at least one of audio sensor data, image sensor data, and biometric sensor data.

5. 10. The computing system of claim 1, wherein the first data comprises pre-existing media content.

6. 6. The computing system of claim 5, wherein the pre-existing media content comprises at least one of audio content, video content, image content, tactile content, or text content.

7. 10. The computing system of claim 1, wherein the first data is received from a network device via at least one of a local area network, a wide area network, or a cellular communication network.

8. 10. The computing system of claim 1, wherein the first data comprises one or more seed parameters.

9. 10. The computing system of claim 1, wherein the network device is a second network device, and the first data is received via the first network device.

10. 10. The computing system of claim 1, wherein the first data comprises data corresponding to text, and the generated synthetic content comprises at least one of audio content, image content, video content, text content, or haptic content.

11. 10. The computing system of claim 1, wherein the second data comprises data related to a non-fungible token (NFT).

12. 12. The computing system of claim 11, wherein the second data comprises data associated with two or more NFTs.

13. 12. The computing system of claim 11, wherein the operations further include: receiving temporary access to the NFT via a network interface.

14. 10. The computing system of claim 1, wherein the second data comprises data related to an output of a smart contract.

15. 10. The computing system of claim 1, wherein the second data comprises data associated with a distributed autonomous organization (DAO).

16. 10. The computing system of claim 1, wherein one or more generative machine learning models have smart contracts.

17. 10. The computing system of claim 1, wherein the first data comprises data received via a blockchain oracle.

18. 10. The computing system of claim 1, wherein the second data comprises data related to access to one or more generative machine learning models.

19. 10. The computing system of claim 1, wherein the network device comprises a server or other computing device that stores a blockchain-based distributed ledger, and the generated synthetic content comprises generated blockchain data.

20. 20. The computing system of claim 19, wherein the generated blockchain data comprises non-fungible tokens (NFTs).

21. 20. The computing system of claim 19, wherein the generated blockchain data comprises a smart contract.

22. 10. The computing system of claim 1, wherein the generated synthetic content comprises at least one of audio content, image content, video content, text content, and haptic content.

23. 10. The computing system of claim 1, wherein the first data comprises a first type of media content and the generated composite content comprises a second type of media content, the second type of media content being different from the first type of media content.

24. 10. The computing system of claim 1, wherein the first data comprises pre-existing media content and the generated composite content comprises a combination of pre-existing media content and new generated content.

25. 2. The computing system of claim 1, wherein the network device comprises a wearable playback device, and the first data comprises data corresponding to a sensor provided on the wearable playback device.

26. 26. The computing system of claim 25, wherein the second data corresponds to data relating to at least one of a user of the wearable playback device, a characteristic of the wearable playback device, or a transaction history associated with the user or the wearable playback device.

27. 10. The computing system of claim 1, wherein the operations further include: storing data corresponding to the first data in a blockchain-based distributed ledger.

28. 10. The computing system of claim 1, wherein the first data comprises data indicative of a user or human presence.

29. 10. The computing system of claim 1, wherein the first data corresponds to a transaction history associated with a network device.

Citation Information

Patent Citations

  • Music content generation

    WO2021163377A1