Generation of digital media based on blockchain data

The media playback system generates and modifies media content in real-time using algorithms and contextual data, addressing the limitations of static content, and integrates with blockchain data for enhanced user experiences.

JP2026053594APending Publication Date: 2026-03-25SONOS INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing media playback systems lack the ability to dynamically generate and modify media content in real-time based on user inputs and contextual data, limiting the potential for personalized and adaptive audio experiences.

Method used

A media playback system that utilizes generative media content, created through algorithms and non-human systems, which can change in real-time based on user inputs, environmental data, and contextual factors, and is associated with blockchain data via non-fungible tokens (NFTs) to enhance user experiences.

Benefits of technology

Enables unique and personalized media experiences that adapt to user emotions, activities, and environments, providing dynamic audio and visual content that enhances user engagement and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053594000001_ABST
    Figure 2026053594000001_ABST
Patent Text Reader

Abstract

This invention provides a method, system, and computer-readable media that enable a unique user experience that cannot be achieved using conventional media playback of pre-recorded content. [Solution] In a coordinator device, the method for playing back generated media content includes the steps of: receiving input parameters; transmitting the input parameters from the coordinator device to a plurality of playback devices, each having a generated media module internally; and transmitting timing data from the coordinator device to the plurality of playback devices so that the playback devices simultaneously play back the generated media content based at least partially on the input parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims the priority of U.S. Application No. 63 / 364,931, filed on May 18, 2022, the entire disclosure of which is incorporated herein by reference.

[0002] This disclosure relates to consumer goods, and more particularly, to methods, systems, products, features, services, and other elements directed to media playback or some aspects thereof.

Background Art

[0003] Until 2002 when SONOS, Inc. began developing a new type of playback system, the option to access and listen to digital audio at high volume settings was limited. Subsequently, Sonos filed one of its first patent applications, titled "Method for Synchronizing Audio Playback between Multiple Networked Devices," in 2003 and began offering its first media playback system for sale in 2005. The Sonos wireless home sound system enables people to experience music from multiple sources via one or more networked playback devices. Through a software - controlled application installed on a controller (e.g., smartphone, tablet, computer, voice - input device), one can play what one desires in any room with a networked playback device. Media content (e.g., songs, podcasts, video sound) can be streamed to the playback devices so that each room with a playback device can play a corresponding different media content. Further, rooms can be grouped together for synchronous playback of the same media content, and / or the same media content can be listened to synchronously in all rooms.

Brief Description of the Drawings

[0004] The features, aspects, and advantages of the technology of this disclosure can be better understood with respect to the following description, the appended claims, and the accompanying drawings, which are listed below. As will be apparent to those skilled in the art, the features shown in the drawings are for illustrative purposes only and variations therein, including different and / or further features and arrangements, are possible. [Figure 1A] Partial break diagram of an environment having a media playback system configured according to the disclosed aspects of the technology. [Figure 1B] Schematic diagram of the media playback system and one or more networks shown in Figure 1A. [Figure 1C] Block diagram of the playback device [Figure 1D] Block diagram of the playback device [Figure 1E] Block diagram of the combined playback device [Figure 1F] Block diagram of a network microphone device [Figure 1G] Block diagram of the playback device [Figure 1H] Partial schematic diagram of the control device [Figure 1I] Schematic diagram of the corresponding media playback system zone [Figure 1J] Schematic diagram of the corresponding media playback system zone [Figure 1K] Schematic diagram of the corresponding media playback system zone [Figure 1L] Schematic diagram of the corresponding media playback system zone [Figure 1M] Schematic diagram of the media playback system area [Figure 2] Functional block diagram of a system for playing generated media content related to this technology. [Figure 3] Functional block diagram of a generation media module relating to an aspect of this technology. [Figure 4] This figure shows an example of an architecture for storing and retrieving generated media content related to an aspect of this technology. [Figure 5]Functional block diagram illustrating data exchange in a system for playing back generated media content according to an embodiment of this technology. [Figure 6] Schematic diagram of an exemplary distributed generation media playback system relating to an embodiment of this technology. [Figure 7] Schematic diagram of another exemplary distributed generation media playback system relating to an aspect of this technology. [Figure 8] Flowchart of the method for playing generated media content related to an aspect of this technology [Figure 9] Flowchart of the method for playing generated media content related to an aspect of this technology [Figure 10] Flowchart of the method for playing generated media content related to an aspect of this technology [Figure 11] Flowchart of the method for playing generated media content related to an aspect of this technology [Figure 12] Flowchart of the method for playing generated media content related to an aspect of this technology [Figure 13] Flowchart of the method for playing generated media content related to an aspect of this technology

[0005] The drawings are for illustrative purposes only; however, as those skilled in the art will see, the art disclosed herein is not limited to the arrangements and / or means shown in the drawings. [Modes for carrying out the invention]

[0006] I. Overview Generative media content is content that is dynamically synthesized, created, and / or modified based on algorithms, whether implemented in software or in a physical model. Generative media content can change over time based solely on algorithms or in relation to contextual data (e.g., user sensor data, environmental sensor data, generational data). In various examples, such generative media content may include generative audio (e.g., music, ambient sounds, etc.), generative visual images (e.g., abstract visual designs that dynamically change lighting, shape, color, etc.), generative scents, generative tactile outputs (e.g., vibrations, tactile outputs), or any other suitable media content or combinations thereof. As described elsewhere in this specification, generative media can be generated at least in part through algorithms and / or non-human systems that utilize rule-based computation to generate novel media content.

[0007] Since generated media content can be dynamically changed in real time, it enables unique user experiences that are not possible with conventional media playback of pre-recorded content. For example, generated audio can be endless and / or dynamic audio that changes as the input to the algorithm changes (e.g., input parameters related to user input, sensor data, media source data, or any other appropriate input data). In some examples, generated audio can be used to direct a user's mood to a desired emotional state using one or more characteristics of the generated audio that change in response to real-time measurements reflecting the user's emotional state. As used in the examples of this technology, the system can provide generated audio based on the user's current and / or desired emotional state, based on the user's activity level, based on the number of users present in the environment, or based on any other appropriate input parameters.

[0008] As another example, generated audio can be created and / or modified based on one or more inputs, such as the user's location or activity, the number of users present in the room, the time of day, or any other input (e.g., determined by one or more sensors or user inputs). For example, when one user is sitting calmly at their desk, the media playback system can automatically generate generated audio content suitable for focused study or work, while when multiple users are present in the room in an excited state with a lot of movement, the same media playback system can automatically generate generated audio suitable for a social gathering or dance party. In various examples, audio characteristics that can be dynamically modified to generate generated audio may include the selection of audio samples or clips, tempo, bass / treble / midrange volume, spatial filtering of the audio output, or any other appropriate audio characteristics. Audio characteristics can be modified by using audio samples that may have different tones or sounds, timing of tones or sounds, and / or desired qualities. In some cases, the playback of the content can also have its characteristics modified by filtering or modulating it, such as equalization, phase, or reverb / delay. During the listening experience, the audio characteristics of the generated music can be modified based on several user inputs, such as time of day, geographical location, weather, or various physiological inputs such as estimated mood, activity level, or heart rate.

[0009] In some cases, generated media content such as a soundscape can be associated with data stored in one or more blockchain databases and / or layers. For example, non-fungible tokens (NFTs) generally consist of data files stored on a blockchain related to digital assets or physical assets (e.g., music, albums, visual artworks, literary works). For example, even though an NFT may reference a digital artwork, it is not common for the NFT itself to contain the digital artwork itself. This is because the required file size is difficult to handle (or too costly) to store in the blockchain layer. Instead, an NFT may consist of metadata (e.g., a URL or other locator) that indicates where the digital artwork is located and / or how to access the artwork. In various examples, NFTs and other data stored via a blockchain or other distributed ledger technology can be used in media content creation. For example, in some examples, an NFT functions as an input to a generative media engine, and as a result, the generated media content has characteristics that are at least partially dependent on the data of a particular NFT or other blockchain.

[0010] Some of the examples described herein can refer to functions performed by a given party such as a "user," "listener," and / or other entity, but it should be understood that this is for illustrative purposes only. The claims should not be construed as requiring acts by the agents of any such examples unless expressly required by the language of the claims themselves.

[0011] In the figures, the same reference numbers generally indicate similar and / or identical elements. To facilitate the description of any particular element, the most significant digit of the reference number refers to the figure in which that element is first introduced. For example, element 110a is first introduced and described with reference to FIG. 1A. Many of the details, dimensions, angles, and other features shown in the figures are merely illustrative of particular examples of the disclosed technology. Thus, other examples can have other details, dimensions, angles, and features without departing from the spirit or scope of the present disclosure. Further, as will be understood by those skilled in the art, additional examples of the various disclosed technologies can be implemented without some of the details described below.

[0012] [[ID=!3]] II. Suitable Operating Environment FIG. 1A is a partial breakaway view of a media playback system 100 distributed within an environment 101 (e.g., a home). The media playback system 100 includes one or more playback devices 110 (individually identified as playback devices 110a - 110n), one or more network microphone devices (``NMDs'') 120 (individually identified as NMDs 120a - 120c), and one or more control devices 130 (individually identified as control devices 130a and 130b).

[0013] As used herein, the term ``playback device'' can generally refer to a network device configured to receive, process, and / or output data of a media playback system. For example, a playback device can be a network device that receives and processes audio content. In some examples, a playback device includes one or more transducers or speakers powered by one or more amplifiers. However, in other examples, a playback device includes neither (or only one) of a speaker and an amplifier. For example, a playback device can include one or more amplifiers configured to drive one or more speakers external to the playback device via a corresponding wire or cable.

[0014] It should be noted that there seems to be an error in the original text where in ID=6, " " is an incomplete or incorrect notation within the sentence. This has been translated as best as possible while maintaining the integrity of the overall text structure.Furthermore, as used herein, the term NMD (i.e., “Network Microphone Device”) can generally refer to a network device configured for audio detection. In some examples, the NMD is a standalone device configured primarily for audio detection. In other examples, the NMD is integrated into a playback device (or vice versa).

[0015] The term "control device" can generally refer to a network device configured to perform functions related to facilitating user access, control, and / or configuration of the media playback system 100.

[0016] Each playback device 110 is configured to receive audio signals or data from one or more media sources (e.g., one or more remote servers or one or more local devices) and to play the received audio signals or data as sound. One or more NMDs 120 are configured to receive voice word commands, and one or more control devices 130 are configured to receive user input. In response to received voice word commands and / or user input, the media playback system 100 can play audio through one or more of the playback devices 110. In certain examples, the playback devices 110 are configured to start playing media content in response to a trigger. For example, one or more of the playback devices 110 may be configured to play a morning playlist when a relevant trigger condition is detected (e.g., the presence of a user in the kitchen, or the operation of a coffee machine). In some examples, the media playback system 100 is configured to play audio from a first playback device (e.g., playback device 110a) in synchronization with a second playback device (e.g., playback device 110b). The interactions between the playback device 110, NMD 120, and / or control device 130 of the media playback system 100 configured according to various examples of this disclosure will be described in more detail below with reference to Figures 1B to 1H.

[0017] In the illustrated example of Figure 1A, environment 101 includes a home having several rooms, spaces, and / or playback zones, including (clockwise from top left) a master bathroom 101a, master bedroom 101b, second bedroom 101c, family room or private room 101d, office 101e, living room 101f, dining room 101g, kitchen 101h, and outdoor patio 101i. Specific examples and illustrations are described below in the context of a home environment, but the technologies described herein may be implemented in other types of environments. In some examples, for example, the media playback system 100 may be implemented in one or more commercial settings (e.g., restaurants, malls, airports, hotels, retail stores or other shops), one or more vehicles (e.g., sports utility vehicles, buses, cars, ships, boats, airplanes), multiple environments (e.g., a combination of a home environment and a vehicle environment), and / or other suitable environments where multi-zone audio may be desirable.

[0018] The media playback system 100 may have one or more playback zones, some of which may correspond to rooms in the environment 101. The media playback system 100 may be established using one or more playback zones after additional zones may be added or removed to form the configuration shown, for example, in Figure 1A. Each zone may be named according to a different room or space, such as office 101e, master bathroom 101a, master bedroom 101b, second bedroom 101c, kitchen 101h, dining room 101g, living room 101f, and / or outdoor patio 101i. In some embodiments, a single playback zone may include multiple rooms or spaces. In certain embodiments, a single room or space may include multiple playback zones.

[0019] In the example illustrated in Figure 1A, the main bathroom 101a, the second bedroom 101c, the office 101e, the living room 101f, the dining room 101g, the kitchen 101h, and the outdoor patio 101i each include one playback device 110, while the main bedroom 101b and the private room 101d include multiple playback devices 110. In the main bedroom 101b, playback devices 110l and 110m may be configured to play audio content synchronously, for example, as individual playback devices 110, as a combined playback zone, as an integrated playback device, and / or any combination thereof. Similarly, in the private room 101d, playback devices 110h to 110j may be configured to play audio content synchronously, for example, as individual playback devices 110, as one or more combined playback devices, and / or as one or more integrated playback devices. Further details regarding combined and integrated playback devices are described below with respect to Figures 1B and 1E.

[0020] In some embodiments, one or more playback zones within environment 101 may each play different audio content. For example, one user may be grilling in patio 101i and listening to hip-hop music being played by playback device 110c, while another user is preparing food in kitchen 101h and listening to classical music being played by playback device 110b. In another example, playback zones may play the same audio content in sync with other playback zones. For example, a user may be in office 101e and hear playback device 110f playing the same hip-hop music being played by playback device 110c on patio 101i. In some embodiments, playback devices 110c and 110f play hip-hop music in sync such that the user perceives the audio content as playing seamlessly (or at least substantially seamlessly) as they move between different playback zones. Further details regarding audio playback synchronization between playback devices and / or zones can be found, for example, in U.S. Patent No. 8,234,395, entitled "System and Method for Synchronizing Operation Between Multiple Independently Clock-Controlled Digital Data Processing Devices," which is incorporated herein by reference in its entirety.

[0021] a. Appropriate media playback system Figure 1B is a schematic diagram of the media playback system 100 and the cloud network 102. For ease of illustration, certain devices of the media playback system 100 and the cloud network 102 have been omitted from Figure 1B. One or more communication links 103 (hereinafter referred to as "link 103") connect the media playback system 100 and the cloud network 102 in a communicative manner.

[0022] Link 103 may comprise, for example, one or more wired networks, one or more wireless networks, one or more wide area networks (WANs), one or more local area networks (LANs), one or more personal area networks (PANs), one or more telecommunications networks (e.g., one or more GSM networks, Code Division Multiple Access (CDMA) networks, Long-Term Evolution (LTE) networks, 5G communication networks, and / or other suitable data transmission protocol networks). Cloud network 102 is configured to deliver media content (e.g., audio content, video content, photos, social media content) to media playback system 100 in response to requests transmitted from media playback system 100 via link 103. In some examples, cloud network 102 is further configured to receive data (e.g., voice input data) from media playback system 100 and, in response, transmit commands and / or media content to media playback system 100.

[0023] The cloud network 102 comprises computing devices 106 (identified separately as a first computing device 106a, a second computing device 106b, and a third computing device 106c). The computing devices 106 may comprise individual computers or servers, such as media streaming service servers, voice service servers, social media servers, and media playback system control servers that store audio and / or other media content. In some examples, one or more of the computing devices 106 comprise a module of a single computer or server. In certain examples, one or more of the computing devices 106 comprise one or more modules, computers, and / or servers. Furthermore, although the cloud network 102 is described above in the context of a single cloud network, in some examples the cloud network 102 comprises multiple cloud networks, including computing devices that are communicatively coupled. Furthermore, although the cloud network 102 is shown in Figure 1B as having three of the computing devices 106, in some examples the cloud network 102 comprises fewer (or more) three computing devices 106.

[0024] The media playback system 100 is configured to receive media content from the network 102 via link 103. The received media content may include, for example, a Uniform Resource Identifier (URI) and / or a Uniform Resource Locator (URL). For example, in some examples, the media playback system 100 may stream, download, or retrieve data from the URI or URL corresponding to the received media content. The network 104 communicatively connects link 103 to at least some of the devices of the media playback system 100 (e.g., one or more of the playback device 110, NMD 120, and / or control device 130). The network 104 may include, for example, a wireless network (e.g., a WiFi network, Bluetooth, Z-Wave network, ZigBee, and / or other suitable wireless communication protocol network) and / or a wired network (e.g., a network including Ethernet, Universal Serial Bus (USB), and / or another suitable wired communication). As will be understood by those skilled in the art, as used herein, “WiFi” can refer to several different communication protocols, including, for example, 2.4 gigahertz (GHz), 5 GHz, and / or other suitable frequencies, including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.11ad, 802.11af, 802.11ah, 802.11ai, 802.11aj, 802.11aq, 802.11ax, 802.11ay, 802.15, etc.

[0025] In some examples, network 104 comprises a dedicated communication network used by the media playback system 100 to send messages between individual devices and / or to and from media content sources (e.g., one or more of the computing devices 106). In certain examples, network 104 is configured to be accessible only to devices within the media playback system 100, thereby reducing interference and conflict with other home devices. However, in other examples, network 104 comprises an existing home communication network (e.g., a home WiFi network). In some examples, link 103 and network 104 comprise one or more of the same networks. In some embodiments, for example, link 103 and network 104 comprise a telecommunications network (e.g., an LTE network, a 5G network). Furthermore, in some examples, the media playback system 100 is implemented without network 104, and devices comprising the media playback system 100 can communicate with each other, for example, via one or more direct connections, PANs, telecommunications networks, and / or other suitable communication links.

[0026] In some examples, audio content sources may be periodically added to or removed from the media playback system 100. In some examples, for instance, the media playback system 100 performs indexing of media items when one or more media content sources are updated, added, and / or removed from the media playback system 100. The media playback system 100 can scan some or all folders and / or directories accessible to the playback device 110 for identifiable media items and generate or update a media content database containing metadata (e.g., title, artist, album, track length) and other relevant information (e.g., URI, URL) for each identifiable media item found. In some examples, for instance, the media content database is stored in one or more of the playback device 110, NMD 120, and / or control devices 130.

[0027] In the illustrated example of Figure 1B, playback devices 110l and 110m comprise group 107a. Playback devices 110l and 110m are located in different rooms within the home and can be temporarily or permanently grouped together in group 107a based on user input received by control device 130a and / or another control device 130 within the media playback system 100. Once placed in group 107a, playback devices 110l and 110m can be configured to play the same or similar audio content synchronously from one or more audio content sources. In certain examples, for instance, group 107a includes a combined zone where playback devices 110l and 110m each contain the left and right audio channels of multi-channel audio content, thereby creating or enhancing the stereo effect of the audio content. In some examples, group 107a includes a further playback device 110. However, in other examples, the media playback system 100 omits group 107a and / or other grouped arrangements of playback devices 110.

[0028] The media playback system 100 includes NMDs 120a and 120d, each having one or more microphones configured to receive voice utterances from a user. In the illustrated example of Figure 1B, NMD 120a is a standalone device, and NMD 120d is integrated into the playback device 110n. NMD 120a is configured to receive, for example, voice input 121 from user 123. In some examples, NMD 120a transmits data related to the received voice input 121 to a Voice Assistant Service (VAS) configured to (i) process the received voice input data and (ii) send corresponding commands to the media playback system 100. In some embodiments, for example, computing device 106c comprises one or more modules and / or servers of a VAS (e.g., SONOS®, AMAZON®, Google®, APPLE®, MICROSOFT®). Computing device 106c can receive voice input data from NMD 120a via network 104 and link 103. In response to receiving the voice input data, computing device 106c processes the voice input data (i.e., "Play Hey Jude by The Beatles") and determines that the processed voice input includes a command to play a song (e.g., "Hey Jude"). Therefore, computing device 106c sends a command to media playback system 100 to play "Hey Jude" by The Beatles from one or more appropriate media services among playback devices 110 (e.g., via one or more computing devices 106).

[0029] b. Appropriate playback device Figure 1C is a block diagram of a playback device 110a having input / output 111. The input / output 111 may include analog I / O 111a (e.g., one or more wires, cables, and / or other suitable communication links configured to carry analog signals) and / or digital I / O 111b (e.g., one or more wires, cables, or other suitable communication links configured to carry digital signals). In some examples, analog I / O 111a is an audio line input connection, including, for example, an auto-sensing 3.5mm audio line input connection. In some examples, digital I / O 111b includes a Sony / Philips Digital Interface Format (S / PDIF) communication interface and / or cable and / or Toshiba Link (TOSLINK) cable. In some examples, digital I / O 111b includes a High Resolution Multimedia Interface (HDMI®) interface and / or cable. In some examples, the digital I / O 111b includes one or more wireless communication links, such as radio frequency (RF), infrared, WiFi, Bluetooth, or another suitable communication protocol. In certain examples, the analog I / O 111a and digital 111b include interfaces (e.g., ports, plugs, jacks) configured to accept connectors for cables transmitting analog and digital signals, respectively, without necessarily including cables.

[0030] The playback device 110a can receive media content (e.g., audio content including music and / or other sounds) from the local audio source 105 via, for example, an input / output 111 (e.g., a cable, wire, PAN, Bluetooth connection, ad-hoc wired or wireless communication network, and / or another suitable communication link). The local audio source 105 may include, for example, a mobile device (e.g., a smartphone, tablet, laptop computer) or another suitable audio component (e.g., a television, desktop computer, amplifier, phonograph, Blu-ray player, memory for storing digital media files). In some embodiments, the local audio source 105 includes a local music library on a smartphone, computer, network-attached storage (NAS), and / or another suitable device configured to store media files. In certain examples, one or more of the playback device 110, NMD 120, and / or control device 130 include the local audio source 105. However, in other examples, the media playback system completely omits the local audio source 105. In some examples, the playback device 110a does not include the input / output 111 and receives all audio content via the network 104.

[0031] The playback device 110a further comprises an electronic device 112, a user interface 113 (e.g., one or more buttons, knobs, dials, touch-sensitive surfaces, displays, touchscreens), and one or more transducers 114 (hereinafter referred to as "transducers 114"). The electronic device 112 is configured to receive audio from an audio source (e.g., a local audio source 105) via an input / output 111, from one or more computing devices 106a to 106c via a network 104 (Figure 1B), amplify the received audio, and output the amplified audio for playback via one or more transducers 114. In some examples, the playback device 110a optionally includes one or more microphones 115 (e.g., a single microphone, multiple microphones, a microphone array) (hereinafter referred to as "microphones 115"). In a specific example, for instance, a playback device 110a having one or more of the optional microphones 115 can operate as an NMD configured to receive audio input from a user and to perform one or more corresponding actions based on the received audio input.

[0032] In the example illustrated in Figure 1C, the electronic device 112 comprises one or more processors 112a (hereinafter referred to as "processor 112a"), memory 112b, software component 112c, network interface 112d, one or more audio processing components 112g (hereinafter referred to as "audio component 112g"), one or more audio amplifiers 112h (hereinafter referred to as "amplifier 112h"), and power supply 112i (e.g., one or more power supplies, power cables, power receptacles, batteries, induction coils, power over Ethernet (PoE) interfaces, and / or other suitable power sources). In some embodiments, the electronic device 112 optionally includes one or more other components 112j (e.g., one or more sensors, video displays, touchscreens, battery charging bases).

[0033] The processor 112a may include a clock-driven computing component configured to process data, and the memory 112b may include a computer-readable medium (e.g., a tangible, non-temporary computer-readable medium, a data storage device, on which one or more software components 112c are loaded) configured to store instructions for performing various operations and / or functions. The processor 112a is configured to execute instructions stored in the memory 112b to perform one or more operations. This operation may include, for example, causing the playback device 110a to retrieve audio data from an audio source (e.g., one or more computing devices 106a-106c (Figure 1B)) and / or another of the playback devices 110. In some examples, the operation may further include causing the playback device 110a to send audio data to another device of the playback devices 110a and / or another device (e.g., one of the NMD 120). A specific example involves pairing playback device 110a with another of one or more playback devices 110 to enable a multi-channel audio environment (e.g., stereo pairs, combined zones).

[0034] The processor 112a may be further configured to perform an operation to synchronize the playback of audio content on the playback device 110a with that of another of the one or more playback devices 110. As will be understood by those skilled in the art, during synchronized playback of audio content on multiple playback devices, the listener is preferably unable to perceive any time delay difference between the playback of audio content on playback device 110a and the playback of audio content on one or more other playback devices 110. Further details regarding audio playback synchronization between playback devices can be found, for example, in U.S. Patent No. 8,234,395, which is incorporated above by reference.

[0035] In some examples, memory 112b is further configured to store data associated with playback device 110a, such as one or more zones and / or zone groups to which playback device 110a is a member, audio sources accessible to playback device 110a, and / or playback queues to which playback device 110a (and / or other playback devices among the one or more playback devices) may be associated. The stored data may include one or more state variables that are periodically updated and used to describe the state of playback device 110a. Memory 112b may also include data associated with the state of one or more other devices of the media playback system 100 (e.g., playback device 110, NMD 120, control device 130). In some embodiments, for example, state data is shared among at least some of the devices of the media playback system 100 at predetermined time intervals (e.g., every 5 seconds, every 10 seconds, every 60 seconds) so that one or more of the devices have the most up-to-date data associated with the media playback system 100.

[0036] The network interface 112d is configured to facilitate the transmission of data between the playback device 110a and one or more other devices on a data network, such as link 103 and / or network 104 (Figure 1B). The network interface 112d is configured to send and receive data corresponding to media content (e.g., audio content, video content, text, photographs) and other signals (e.g., non-transient signals), including digital packet data with Internet Protocol (IP)-based source and / or IP-based destination addresses. The network interface 112d can parse the digital packet data so that the electronic device 112 can properly receive and process data destined for the playback device 110a.

[0037] In the illustrated example in Figure 1C, the network interface 112d comprises one or more wireless interfaces 112e (hereinafter referred to as "wireless interfaces 112e"). The wireless interfaces 112e (e.g., appropriate interfaces including one or more antennas) can be configured to wirelessly communicate with one or more other devices (e.g., one or more of other playback devices 110, NMD 120, and / or control devices 130) that are communicably coupled to the network 104 (Figure 1B) according to an appropriate wireless communication protocol (e.g., WiFi, Bluetooth, LTE). In some examples, the network interface 112d optionally includes a wired interface 112f (e.g., an interface or receptacle configured to receive network cables such as Ethernet, USB-A, USB-C, and / or Thunderbolt cables) configured to communicate with other devices via a wired connection according to an appropriate wired communication protocol. In certain examples, the network interface 112d includes a wired interface 112f and excludes the wireless interfaces 112e. In some examples, the electronic device 112 completely excludes the network interface 112d and sends and receives media content and / or other data via a separate communication path (e.g., input / output 111).

[0038] The audio component 112g is configured to process and / or filter data, including media content, received by the electronic device 112 (for example, via the input / output 111 and / or network interface 112d) to generate an output audio signal. In some examples, the audio processing component 112g comprises, for example, one or more digital-to-analog converters (DACs), audio preprocessing components, audio enhancement components, digital signal processors (DSPs), and / or other suitable audio processing components, modules, circuits, etc. In certain examples, one or more of the audio processing components 112g may comprise one or more subcomponents of the processor 112a. In some examples, the electronic device 112 omits the audio processing component 112g. In some embodiments, for example, the processor 112a executes instructions stored in memory 112b to perform audio processing operations to generate an output audio signal.

[0039] Amplifier 112h is configured to receive and amplify the audio output signal generated by the audio processing component 112g and / or processor 112a. Amplifier 112h may comprise electronic devices and / or components configured to amplify the audio signal to a level sufficient to drive one or more of the transducers 114. In some examples, for instance, amplifier 112h includes one or more switching or Class D power amplifiers. However, in other examples, the amplifier includes one or more other types of power amplifiers (e.g., linear gain power amplifiers, Class A amplifiers, Class B amplifiers, Class AB amplifiers, Class C amplifiers, Class D amplifiers, Class E amplifiers, Class F amplifiers, Class G amplifiers and / or Class H amplifiers, and / or other suitable types of power amplifiers). In certain examples, amplifier 112h comprises two or more suitable combinations of the aforementioned types of power amplifiers. Furthermore, in some examples, individual components of amplifier 112h correspond to individual components of transducer 114. However, in other examples, the electronic device 112 includes one of several amplifiers 112h configured to output the amplified audio signal to multiple transducers 114. In some other examples, the electronic device 112 omits amplifier 112h.

[0040] The transducer 114 (e.g., one or more speakers and / or speaker drivers) receives the amplified audio signal from the amplifier 112h and renders or outputs the amplified audio signal as sound (e.g., an audible sound wave having a frequency of approximately 20 Hz to 20 kHz). In some examples, the transducer 114 may comprise a single transducer. However, in other examples, the transducer 114 may comprise multiple audio transducers. In some examples, the transducer 114 may comprise multiple types of transducers. For example, the transducer 114 may include one or more low-frequency transducers (e.g., subwoofers, woofers), mid-frequency transducers (e.g., midrange transducers, midwoofers), and one or more high-frequency transducers (e.g., one or more tweeters). As used herein, “low frequency” can generally refer to audible frequencies below approximately 500 Hz, “mid frequency” can generally refer to audible frequencies between approximately 500 Hz and approximately 2 kHz, and “high frequency” can generally refer to audible frequencies above 2 kHz. However, in certain examples, one or more of the transducers 114 may be transducers that do not adhere to the aforementioned frequency ranges. For example, one of the transducers 114 may be an intermediate woofer transducer configured to output sound at frequencies between approximately 200 Hz and approximately 5 kHz.

[0041] As an example, SONOS, Inc. currently offers (or is offering) certain playback devices for sale, including, for example, "SONOS ONE," "PLAY:1," "PLAY:3," "PLAY:5," "PLAYBAR," "PLAYBASE," "CONNECT:AMP," "CONNECT," and "SUB." In addition to or instead of these, other suitable playback devices may be used to implement the example playback devices disclosed herein. Furthermore, as will be apparent to those skilled in the art, playback devices are not limited to the examples described herein or the SONOS product offerings. In some examples, for example, one or more playback devices 110 comprise wired or wireless headphones (e.g., over-ear headphones, in-ear headphones, earphones). In other examples, one or more of the playback devices 110 comprise a docking station and / or interface configured to interact with a docking station for personal mobile media playback devices. In certain examples, the playback device may be integrated with another device or component, such as a television, lighting fixture, or any other device for indoor or outdoor use. In some examples, the playback device omits a user interface and / or one or more transducers. For example, Figure 1D is a block diagram of a playback device 110p having inputs / outputs 111 and electronic equipment 112, but without a user interface 113 or transducer 114.

[0042] Figure 1E is a block diagram of a coupled playback device 110q comprising a playback device 110i (e.g., a subwoofer) (Figure 1A) and an ultrasonically coupled playback device 110a (Figure 1C). In the illustrated example, playback devices 110a and 110i are separate playback devices 110 housed in separate enclosures. However, in some examples, the coupled playback device 110q comprises a single enclosure housing both playback devices 110a and 110i. The coupled playback device 110q can be configured to process and reproduce sound in a different manner than uncoupled playback devices (e.g., playback device 110a in Figure 1C) and / or paired or coupled playback devices (e.g., playback devices 110l and 110m in Figure 1B). In some examples, for instance, playback device 110a is a full-range playback device configured to render low-frequency, mid-frequency, and high-frequency audio content, and playback device 110i is a subwoofer configured to render low-frequency audio content. In some embodiments, playback device 110a, when coupled with a first playback device, is configured to render only the mid-range and high-frequency components of a particular audio content, while playback device 110i renders only the low-frequency components of the particular audio content. In some examples, coupled playback device 110q includes additional playback devices and / or another coupled playback device.

[0043] c. A suitable network microphone device (NMD) Figure 1F is a block diagram of the NMD120a (Figures 1A and 1B). The NMD120a includes one or more audio processing components 124 (hereinafter, "audio components 124") and several components described in relation to the playback device 110a (Figure 1C), which includes a processor 112a, memory 112b, and microphone 115. The NMD120a optionally includes other components also included in the playback device 110a (Figure 1C), such as a user interface 113 and / or transducer 114. In some examples, the NMD120a is configured as a media playback device (e.g., one or more of the playback devices 110) and further includes, for example, an audio component 112g (Figure 1C), an amplifier 114, and / or one or more other playback device components. In certain examples, the NMD120a includes Internet of Things (IoT) devices such as a thermostat, alarm panel, fire and / or smoke detector. In some examples, the NMD120a comprises a microphone 115, an audio processor 124, and only some of the components of the electronic device 112 described with respect to Figure 1B. In some embodiments, for example, the NMD120a includes a processor 112a and memory 112b (Figure 1B), while omitting one or more other components of the electronic device 112. In some examples, the NMD120a includes additional components (e.g., one or more sensors, a camera, a thermometer, a barometer, a hygrometer).

[0044] In some examples, the NMD can be incorporated into the playback device. Figure 1G is a block diagram of a playback device 110r equipped with an NMD 120d. The playback device 110r may comprise many or all of the components of the playback device 110a and may further include a microphone 115 and an audio processor 124 (Figure 1F). The playback device 110r may have an integrated control device 130c. The control device 130c may comprise a user interface (e.g., user interface 113 in Figure 1B) configured to receive user input (e.g., touch input, voice input) without involving a separate control device. However, in other examples, the playback device 110r receives commands from another control device (e.g., control device 130a in Figure 1B).

[0045] Referring again to Figure 1F, the microphone 115 is configured to acquire, capture, and / or receive sound from the environment in which the NMD 120a is located (e.g., environment 101 in Figure 1A) and / or the room. The received sound may include, for example, speech utterances, audio played by the NMD 120a and / or another playback device, background noise, ambient noise, etc. The microphone 115 converts the received sound into an electrical signal to generate microphone data. The voice processing unit 124 receives and analyzes the microphone data to determine whether there is voice input in the microphone data. Voice input may include, for example, an activation word followed by an utterance containing a user request. As will be understood by those skilled in the art, an activation word is a word or other audio cue that signifies user voice input. For example, when querying an Amazon® VAS, a user may utter the activation word "Alexa". Other examples include "Okay, Google" to invoke a Google® VAS and "Hey, Siri" to invoke an Apple® VAS.

[0046] After detecting an activation word, the voice processor 124 monitors microphone data for user requests associated with the voice input. User requests may include commands to control third-party devices such as a thermostat (e.g., NEST® thermostat), a lighting device (e.g., PHILIPS HUE® lighting device), or a media playback device (e.g., Sonos® playback device). For example, a user can set the temperature in their home by uttering the activation word "Alexa" (e.g., environment 101 in Figure 1A) followed by the utterance "Please set the thermostat to 68 degrees." The user can then turn on the lighting device in the living room area of ​​their home by uttering the same activation word followed by the utterance "Turn on the living room." The user may similarly speak an activation word followed by a request to play a specific song, album, or music playlist on a playback device in their home.

[0047] d. Appropriate control device Figure 1H is a partial schematic diagram of the control device 130a (Figures 1A and 1B). As used herein, the term “control device” can be used interchangeably with “controller” or “control system.” Among other features, the control device 130a is configured to receive user input related to the media playback system 100 and, in response, cause one or more devices within the media playback system 100 to perform an action or operation corresponding to the user input. In the illustrated example, the control device 130a comprises a smartphone (e.g., iPhone®, Android phone) with media playback system controller application software installed. In some examples, the control device 130a comprises, for example, a tablet (e.g., iPad®), a computer (e.g., laptop computer, desktop computer), and / or another suitable device (e.g., television, car audio head unit, IoT device). In a particular example, the control device 130a comprises a dedicated controller for the media playback system 100. In another example, as previously mentioned with respect to Figure 1G, the control device 130a is incorporated into another device within the media playback system 100 (e.g., one or more of the playback device 110, NMD 120, and / or other suitable devices configured to communicate over a network).

[0048] The control device 130a includes an electronic device 132, a user interface 133, one or more speakers 134, and one or more microphones 135. The electronic device 132 comprises one or more processors 132a (hereinafter referred to as "processor 132a"), a memory 132b, a software component 132c, and a network interface 132d. The processor 132a can be configured to perform functions related to facilitating user access, control, and configuration of the media playback system 100. The memory 132b can be a data storage device that can load one or more software components executable by the processor 112a to perform these functions. The software component 132c can be an application and / or other executable software configured to facilitate control of the media playback system 100. The memory 112b can be configured to store, for example, the software component 132c, media playback system controller application software, and / or other data related to the media playback system 100 and the user.

[0049] The network interface 132d is configured to facilitate network communication between the control device 130a and one or more other devices in the media playback system 100, and / or one or more remote devices. In some examples, the network interface 132d is configured to operate according to one or more appropriate communication industry standards (e.g., infrared, wireless, wired standards including IEEE 802.3, wireless standards including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G, LTE). The network interface 132d can be configured to transmit and / or receive data from, for example, the playback device 110, the NMD 120, other control devices 130, one of the computing devices 106 in Figure 1B, one or more other devices with media playback systems, etc. The transmitted and / or received data may include, for example, playback device control commands, state variables, playback zones, and / or zone group configurations. For example, based on user input received in user interface 133, network interface 132d can send playback device control commands (e.g., volume control, audio playback control, audio content selection) from control device 130 to one or more of the playback devices 110. Network interface 132d can also send and / or receive configuration changes, such as, among other things, adding / removing one or more playback devices 110 to / from a zone, adding / removing one or more zones to / from a zone group, forming combined or integrated players, and separating one or more playback devices from combined or integrated players. Descriptions of adding zones and groups can be found below with respect to Figures 1I to 1M.

[0050] The user interface 133 is configured to receive user input and facilitate control of the media playback system 100. The user interface 133 includes media content technology 133a (e.g., album art, lyrics, video), a playback status indicator 133b (e.g., elapsed time and / or remaining time indicator), a media content information area 133c, a playback control area 133d, and a zone indicator 133e. The media content information area 133c may include the display of relevant information (e.g., title, artist, album, genre, release year) about the media content currently being played and / or media content in the queue or playlist. The playback control area 133d may include selectable icons (e.g., via touch input and / or via a cursor or another appropriate selector) for performing playback actions on one or more playback devices in a selected playback zone or zone group, such as play or pause, fast forward, rewind, skip to next, skip to previous, enter / exit shuffle mode, enter / exit repeat mode, enter / exit crossfade mode, etc. The playback control area 133d may also include selectable icons for changing equalization settings, playback volume, and / or other appropriate playback operations. In the illustrated example, the user interface 133 comprises a display presented on the touchscreen interface of a smartphone (e.g., iPhone®, Android phone). However, in some examples, user interfaces for various formats, styles, and interactive sequences can be alternatively implemented on one or more network devices to provide equivalent control access to the media playback system.

[0051] One or more speakers 134 (e.g., one or more transducers) may be configured to output sound to the user of the control device 130a. In some examples, one or more speakers comprise individual transducers configured to output low frequencies, mid frequencies, and / or high frequencies, respectively. In some embodiments, for example, the control device 130a is configured as a playback device (e.g., one of the playback devices 110). Similarly, in some examples, the control device 130a is configured as an NMD (e.g., one of the NMDs 120) and receives voice commands and other sounds via one or more microphones 135.

[0052] One or more microphones 135 may comprise, for example, one or more condenser microphones, electret condenser microphones, dynamic microphones, and / or other suitable types of microphones or transducers. In some examples, two or more of the microphones 135 are arranged to capture location information of an audio source (e.g., speech, audible sound) and / or are configured to facilitate filtering of background noise. Furthermore, in certain examples, the control device 130a is configured to operate as a playback device and NMD. However, in other examples, the control device 130a omits one or more speakers 134 and / or one or more microphones 135. For example, the control device 130a may comprise a part of the electronic equipment 132 and a user interface 133 (e.g., a touchscreen) without speakers or microphones (e.g., a thermostat, an IoT device, a network device).

[0053] Appropriate playback device configuration Figures 1I to 1M show exemplary configurations of playback devices in zones and zone groups. Referring first to Figure 1M, in one example, a single playback device may belong to a zone. For example, playback device 110g in the second bedroom 101c (Figure 1A) may belong to zone C. In some implementations described below, multiple playback devices can be “combined” to form “combined pairs,” which together form a single zone. For example, playback device 110l (e.g., left playback device) can be combined with playback device 110l (e.g., left playback device) to form zone A. Combined playback devices may have different playback responsibilities (e.g., channel responsibilities). In another embodiment described below, multiple playback devices can be merged to form a single zone. For example, playback device 110h (e.g., front playback device) may be merged with playback device 110i (e.g., subwoofer) and playback devices 110j and 110k (e.g., left and right surround speakers, respectively) to form a single zone D. In another example, playback devices 110g and 110h can be merged to form a merged group or zone group 108b. The merged playback devices 110g and 110h do not necessarily have to be assigned different playback responsibilities. That is, the merged playback devices 110h and 110i can each play audio content in the same way as if they were not merged, separate from playing audio content in synchronously.

[0054] Each zone within the media playback system 100 may be provided for control as a single user interface (UI) entity. For example, Zone A may be provided as a single entity called the main bathroom. Zone B may be provided as a single entity called the master bedroom. Zone C may be provided as a single entity called the second bedroom.

[0055] The coupled playback devices may have different playback responsibilities, such as responsibility for a specific audio channel. For example, as shown in Figure 1-I, playback devices 110l and 110m may be coupled to produce or enhance the stereo effect of audio content. In this example, playback device 110l may be configured to play the left channel audio component, while playback device 110k may be configured to play the right channel audio component. In some implementations, such stereo coupling is sometimes referred to as "pairing".

[0056] Furthermore, the coupled playback devices may have additional and / or different respective speaker drivers. As shown in Figure 1J, playback device 110h, named Front, may be coupled with playback device 110i, named SUB. Front device 110h can be configured to render the mid-range to high-range frequencies, and SUB device 110i can be configured to render the low frequencies. However, if not coupled, Front device 110h can be configured to render the entire frequency range. As another example, Figure 1K shows Front device 110h and SUB device 110i further coupled with left playback device 110j and right playback device 110k, respectively. In some implementations, the right device 110j and left device 110k can be configured to form a surround or "satellite" channel in a home theater system. The coupled playback devices 110h, 110i, 110j, and 110k can form a single zone D (Figure 1M).

[0057] Merged playback devices do not need to be assigned playback responsibilities, and each playback device can render the full range of audio content it is capable of. Nevertheless, merged devices may be represented as a single UI entity (i.e., a zone, as described above). For example, playback devices 110a and 110n in the main bathroom have a single UI entity in zone A. In one example, playback devices 110a and 110n can each output the full range of audio content that each playback device 110a and 110n can synchronously produce.

[0058] In some examples, an NMD is coupled or merged with another device to form a zone. For example, NMD120b may be coupled with playback device 110e to form a zone F called the Living Room. In other examples, a standalone network microphone device may be in a zone itself. However, in other examples, a standalone network microphone device may not be associated with a zone. Further details regarding associating a network microphone device and a playback device as designated or default devices can be found, for example, in U.S. Patent Application No. 15 / 438,749, which was referenced earlier.

[0059] The zones of individual devices, combined devices, and / or merged devices can be grouped to form zone groups. For example, referring to Figure 1M, zone A can be grouped with zone B to form zone group 108a containing the two zones. Similarly, zone G may be grouped with zone H to form zone group 108b. As another example, zone A may be grouped with one or more other zones CI. Zones A-I can be grouped and ungrouped in numerous ways. For example, three, four, five, or more (e.g., all) of zones A-I may be grouped. Once grouped, the zones of individual and / or combined playback devices can play audio synchronously with each other, as described in U.S. Patent No. 8,234,395, which was referenced earlier. Playback devices may be dynamically grouped and ungrouped to form new or different groups that play audio content synchronously.

[0060] In various embodiments, zones within an environment may be the default names of zones within a group or a combination of names of zones within a zone group. For example, zone group 108b may be assigned a name such as "Dining + Kitchen," as shown in Figure 1M. In some examples, zone groups may be given unique names selected by the user.

[0061] Certain data may be stored in the memory of a playback device (e.g., memory 112b in Figure 1C) as one or more state variables that are periodically updated and used to describe the state of a playback zone, a playback device, and / or a group of zones associated with a playback zone. The memory may also include data that is shared from time to time between devices so that it is associated with the state of other devices in the media system and one or more of the devices have the most up-to-date data associated with the system.

[0062] In some examples, memory can store instances of various variable types related to state. Variable instances may be stored with identifiers (e.g., tags) corresponding to their type. For example, a particular identifier might be a first type "a1" for identifying a zone's regeneration device, a second type "b1" for identifying regeneration devices that can be combined within a zone, and a third type "c1" for identifying a zone group to which a zone can belong. As a relevant example, an identifier associated with the second bedroom 101c might indicate that the regeneration device is the only regeneration device in zone C and not within a zone group. An identifier associated with Den might indicate that Den is not grouped with other zones but contains combined regeneration devices 110h-110k. An identifier associated with the dining room might indicate that the dining room is part of the dining + kitchen zone group 108b and that devices 110b and 110d are grouped together (Figure 1L). An identifier associated with the kitchen might indicate the same or similar information by indicating that the kitchen is part of the dining + kitchen zone group 108b. Other exemplary zone variables and identifiers are described below.

[0063] In yet another example, the media playback system 100 may be a variable or identifier representing other associations of zones and zone groups, such as an identifier associated with an area, as shown in Figure 1M. An area may include clusters of zone groups and / or zones that are not within a zone group. For example, Figure 1M shows an upper area 109a containing zones A-D and a lower area 109b containing zones E-I. In one embodiment, an area may be used to refer to a zone group and / or cluster of zones that share one or more zones and / or zone groups of another cluster. In another embodiment, this is different from a zone group that does not share a zone with another zone group. Further examples of techniques for implementing areas can be found, for example, in U.S. Patent Application No. 15 / 682,506, filed on 21 August 2017 and titled “Name-Based Room Associations,” and U.S. Patent No. 8,483,853, filed on 11 September 2007 and titled “Control and Manipulation of Grouping in Multizone Media Systems.” Each of these applications is incorporated herein by reference in its entirety. In some cases, the media playback system 100 does not need to implement areas, in which case the system does not need to store variables associated with areas.

[0064] III. Multi-device playback of generated media content Figure 2 is a functional block diagram of system 200 for playback of generated media content. As previously stated, generated media content may include any media content (e.g., audio, video, audiovisual output, tactile output, or any other media content) that is dynamically created, synthesized, and / or modified by a non-human, rule-based process such as an algorithm or model. This creation or modification may be performed for playback in real time or near real time. In addition to or instead of this, generated media content may be generated or modified asynchronously (e.g., before playback is requested), and specific items of the generated media content may then be selected later for playback. As used herein, “generated media module” includes any system that can generate generated media content based on one or more inputs, whether implemented in software, a physical model, or a combination thereof. In some examples, such generated media content may be created as entirely new, or include new media content that can be created by mixing, combining, manipulating, or otherwise modifying one or more existing parts of media content. As used herein, “Generate Media Content Model” includes any set of algorithms, schemas, or rules that can be used to generate new generate media content using one or more inputs (e.g., sensor data, artist-provided parameters, media segments such as audio clips or samples). In the example, a generate media module may use different generate media content models to generate various kinds of generate media content. In some cases, an artist or other collaborator may interact with a generate media content model, create it, and / or update it to generate specific kinds of generate media content.While some examples throughout this explanation refer to audio content, the principles disclosed herein can, in some examples, be applied to other types of media content, such as video, audiovisual, haptic, or others.

[0065] As shown in Figure 2, the system 200 includes a generating media group coordinator 210 that communicates with generating media group members 250a and 250b, as well as a sensor data source 218, a media content source 220, and a control device 130. Such communication can be performed over a network 102 that can include any suitable wired or wireless network connection or a combination thereof (e.g., a WiFi network, Bluetooth, Z-Wave network, ZigBee, Ethernet connection, Universal Serial Bus (USB) connection, etc.), as previously described.

[0066] One or more remote computing devices 106 can also communicate with the group coordinator 210 and / or group members 250a and 250b via the network 102. In various examples, the remote computing device 106 may be a cloud-based server associated with a device manufacturer, a media content provider, a voice assistant service, or other appropriate entity. As shown in Figure 2, the remote computing device 106 may include a generated media module 214. As will be described in more detail elsewhere in this specification, the remote computing device 106 can generate generated media content remotely from local devices (e.g., the coordinator 210 and members 250a and 250b). The generated media content can then be transmitted to one or more local devices for playback. In addition to or instead of this, the generated media content can be generated entirely or partially via local devices (e.g., the group coordinator 210 and / or group members 250a and 250b). In some examples, the group coordinator 210 can itself be a remote computing device, communicating with group members 250a and 250b via a wide area network, and the devices do not need to be located in the same place within the same environment (e.g., home, office, etc.).

[0067] a. Example of generated media group operation In the illustrated example, the generated media group includes a generated media group coordinator 210 (also referred to herein as the “coordinator device 210”) and first and second generated media group members 250a and 250b (also collectively referred to herein as “first member device 250a”, “second member device 250b”, and “member device 250”). Optionally, one or more remote computing devices 106 may also form part of the generated media group. During operation, these devices can communicate with each other and / or with other components (e.g., a sensor data source 218, a control device 130, a media content source 220, or any other suitable data source or component) to facilitate the generation and playback of generated media content.

[0068] In various examples, some or all of devices 210 and / or 250 may be located in the same place within the same environment (e.g., the same home, store, etc.). In some examples, at least some of devices 210 and / or 250 may be located separately from each other, for example, in different homes, different cities, etc.

[0069] The coordinator device 210 and / or member device 250 may include some or all of the components of the playback device 110 or network microphone device 120 described above with respect to Figures 1A to 1H. For example, the coordinator device 210 and / or member device 250 may optionally include playback components 212 (e.g., transducers, amplifiers, audio processing components, etc.), or such components may be omitted in some cases.

[0070] In some examples, the coordinator device 210 is the playback device itself and can therefore also operate as a member device 250. In other examples, the coordinator device 210 can be connected to one or more member devices 250 (e.g., via a direct wired connection or via the network 102), but the coordinator device 210 does not play the generated media content itself. In various examples, the coordinator device 210 can be implemented on a bridge device on a local network, a playback device that is not itself part of the generated media group (i.e., the playback device itself does not play the generated media content), and / or on a remote computing device (e.g., a cloud server).

[0071] In various examples, one or more devices may include a generating media module 214 on top of it. Such a generating media module 214 can generate new synthesized media content based on one or more inputs, for example, using an appropriate generating media content model. As shown in Figure 2, in some examples, the coordinator device 210 may include a generating media module 214 for generating generated media content, which can then be transmitted to member devices 250a and 250b for simultaneous and / or synchronous playback. In addition to or instead of this, some or all of the member devices 250 (e.g., member device 250b shown in Figure 2) may include a generating media module 214, which can be used by the member devices 250 to locally generate generated media content based on one or more inputs. In various examples, generated media content can be generated via a remote computing device 106 using one or more input parameters optionally received from a local device. This generated media content can then be transmitted to one or more local devices for coordination and / or playback.

[0072] In some examples, at least some of the member devices 250 do not include a generating media module 214. Alternatively, each member device 250 may include a generating media module 214 and be configured to generate generating media content locally. In at least some examples, none of the member devices 250 include a generating media module 214. In such cases, the generating media content can be generated by the coordinator device 210. Such generated media content can then be transmitted to the member devices 250 for simultaneous and / or synchronous playback.

[0073] In the example shown in Figure 2, the coordinator device 210 further includes a coordination component 216. As will be described in more detail herein, the coordinator device 210 may, in some cases, facilitate playback of generated media content through multiple different playback devices (which may or may not include the coordinator device 210 itself). During operation, the coordination component 216 is configured to facilitate synchronization of both generated media creation (for example, using one or more generated media modules 214 which may be distributed across various devices) and generated media playback. For example, the coordinator device 210 may send timing data to member devices 250 to facilitate synchronized playback. In addition to or instead of the above, the coordinator device 210 may send inputs, generation media model parameters, or other data relating to the generation media module 214 to one or more member devices 250 so that the member devices 250 can generate the generation media locally (for example, using a locally stored generation media module 214), and / or update or modify the generation media module 214 based on inputs received from the coordinator device 210.

[0074] As will be described in more detail elsewhere in this specification, the generated media module 214 can be configured to generate generated media based on one or more inputs using a generated media content model. The inputs may include sensor data (e.g., provided by the sensor data source 218), user input (e.g., received from the control device 130, or via direct user interaction with the coordinator device 210 or member device 250), and / or the media content source 220. For example, the generated media module 214 can generate and continuously modify generated audio by adjusting various characteristics of the generated audio based on one or more input parameters (e.g., sensor data about one or more users to devices 210, 250).

[0075] b. Exemplary media content sources In various examples, the media content source 220 may include one or more local and / or remote media content sources. For example, the media content source 220 may include one or more local audio sources 105 as described above (e.g., audio received via input / output connections from a mobile device (e.g., a smartphone, tablet, laptop computer) or another suitable audio component (e.g., a television, desktop computer, amplifier, phonograph, Blu-ray player, memory for storing digital media files)). In addition to or instead of this, the media content source 220 may include one or more remote computing devices accessible via a network interface (e.g., via communication over network 102). Such remote computing devices may include individual computers or servers, such as media streaming service servers that store audio and / or other media content.

[0076] In various examples, the media available through the media content source 220 may include a complete sound, a song, a part of a song (e.g., a sample), or pre-recorded audio segments in the form of any audio component (e.g., pre-recorded audio of a particular instrument, a synthesized beat or other audio segment, non-musical audio such as spoken language or natural sounds). During operation, such media can be used by the generating media module 214 to generate generated media content by, for example, combining, mixing, overlaying, manipulating, or otherwise modifying the retrieved media content to generate new generated media content for playback through one or more devices. In some examples, the generated media content may take the form of a combination of pre-recorded audio segments (e.g., a pre-recorded song, a recording of spoken language, etc.) and new synthesized audio created and overlaid with the pre-recorded audio. As used herein, “generated media content” or “generated media content” may include any such combination.

[0077] c. Exemplary generating media module As described above, the generating media module 214 can include any system capable of generating generated media content based on one or more inputs, whether instantiated in software, a physical model, or a combination thereof. In various examples, the generating media module 214 can utilize a generating media content model, which can include one or more algorithms or mathematical models that determine how media content is generated based on relevant input parameters. In some cases, the algorithms and / or mathematical models themselves can be updated over time, for example, based on instructions received from one or more remote computing devices (e.g., cloud servers associated with a music service or other entity), or inputs received from other group member devices in the same or different environment, or any other suitable inputs. In some examples, different devices in a group may have different generating media modules 214 on them, for example, using a first member device that has a different generating media module 214 than a second member device. In other cases, each device in a group having a generating media module 214 may include substantially the same model or algorithm.

[0078] Any suitable algorithm or combination of algorithms can be used to generate generated media content. Examples of such algorithms include machine learning techniques (e.g., generative adversarial networks, neural networks, etc.), formal grammars, Markov models, finite-state automata, and / or the use of any algorithm implemented in currently available offerings such as JukeBox by OpenAI, AWS DeepComposer by Amazon, Magenta by Google, and Amper AI by Amper Music. In various examples, the generated media module 214 can utilize any suitable generating algorithm that currently exists or may be developed in the future.

[0079] In accordance with the above description, generating media content (e.g., audio content) may include modifying various characteristics of the media content in real time and / or algorithmically generating new media content in real time or near real time. In the context of audio content, this can be achieved by storing several audio samples in a database (e.g., within the media content source 220) that can be remotely located and made accessible by the coordinator device 210 and / or member device 250 via the network 102, or by maintaining the audio samples locally on devices 210, 250 themselves. An audio sample can be associated with one or more metadata tags corresponding to one or more audio characteristics of the sample. For example, a given sample can be associated with a metadata tag indicating that the sample contains audio of a specific frequency or frequency range (e.g., bass / mid / treble) or a specific instrument, genre, tempo, key, release date, geographical area, timbre, reverb, distortion, sonic texture, or any other audio characteristics that may be revealed.

[0080] During operation, the generating media module 214 (for example, of the coordinator device 210 and / or the second member device 250b) can search for specific audio samples based on their associated tags and mix them to create generated audio. The generated audio can evolve in real time as the generating media module 214 searches for audio samples with different tags and / or different audio samples with the same or similar tags. The audio samples that the generating media module 214 searches for can depend on one or more inputs, such as various user inputs like sensor data, time, geographical location, weather, or mood selection, or physiological inputs like heart rate. In this way, as the input changes, the generated audio also changes. For example, if the user selects a mood input of calming or relaxing, the generating media module 214 can search for audio samples and mix them with tags corresponding to audio content that can calm or relax the user. Examples of such audio samples include audio samples tagged as having a slow tempo or low harmonic complexity, or audio samples that are predetermined to be calm and relaxing and tagged in that way. In some examples, audio samples can be identified as calming or relaxing based on an automated process that analyzes the temporal and spectral content of the signal. Other examples are similarly possible. In any of the examples herein, the generating media module 214 can adjust the characteristics of the generated audio by searching for and mixing audio samples associated with different metadata tags or other appropriate identifiers.

[0081] Modifying the characteristics of generated audio can include manipulating one or more of the following: volume, balance, removal of specific instruments or tones, tempo, gain, reverb, spectral equalization, timbre, or sound texture. In some examples, generated audio can be played differently on different devices, for example, by emphasizing certain characteristics of the generated audio on the specific playback device closest to the user. For instance, the nearest playback device might emphasize a particular instrument, beat, tone, or other characteristic, while the remaining playback devices could function as background audio sources.

[0082] As described elsewhere in this specification, the media content module 214 can be configured to generate media intended to guide the user's mood and / or physiological state in a desired direction. In some examples, the user's current state (e.g., mood, emotional state, activity level, etc.) is constantly and / or repeatedly monitored or measured (e.g., at predetermined intervals) to ensure that the user's current state is transitioning toward a desired state or at least not in the opposite direction. In such examples, the generated audio content can be modified to guide the user's current state toward a desired ending state.

[0083] In any of the examples herein, the generating media module may use hysteresis to avoid rapid adjustments to the generated audio that could negatively impact the listening experience. For example, if the generating media module changes the media based on user position input to the playback device, the playback device may rapidly change the generated audio in any of the methods described herein as the user moves closer to or further away from the playback device. Such abrupt adjustments can be unpleasant for the user. To mitigate these rapid adjustments, the generating media module 214 may be configured to use hysteresis by delaying adjustments to the generated audio for a predetermined period when user movement or other activity triggers an adjustment. For example, if the playback device detects that the user has moved within a threshold distance of the playback device, instead of immediately performing one of the aforementioned adjustments, the playback device may wait for a predetermined amount of time (e.g., a few seconds) before making an adjustment. If the user remains within the threshold distance after the predetermined amount of time, the playback device may proceed with adjusting the generated audio. However, if the user is not within the threshold distance after the predetermined amount of time, the generating media module 214 may refrain from adjusting the generated audio. The generation media module 214 can similarly apply hysteresis to other generation media adjustments described herein.

[0084] Figure 3 shows a flowchart of process 300 for generating generated audio content using various input parameters. In various examples, one or more of these input parameters can be modified based on user input. For example, an artist can select various parameters, constraints, or available audio segments shown in Figure 2, and these selections can at least partially determine the final output of the generated audio content. As previously mentioned, such a generated media module may be stored and operated on one or more playback devices for local playback (e.g., via the same playback device and / or via other playback devices commutably coupled via a local area network). In addition to or instead of this, such a generated media module may be stored and operated on one or more remote computing devices, and the resulting output may be sent to one or more remote devices via a wide area network for playback.

[0085] As illustrated, the process begins in block 302 and proceeds to block 304, where the clock / metronome receives inputs of tempo 306 and time signature 308. The tempo 306 and time signature 308 can be selected by the artist or automatically determined or generated using a model. The process then proceeds to block 310, where a chord change can be triggered, receiving a chord change frequency parameter 312 as input. The artist may choose to have a higher chord change frequency in music intended for a higher energy experience (e.g., dance music, uplifting ambient music, etc.). Conversely, a lower chord change frequency may be associated with a lower energy output (e.g., quiet music).

[0086] In block 314, a code is selected from the available code segments 316. Multiple code information parameters 318, 320, and 322 can also be provided as inputs to the code segment 316. These inputs can be used to determine a specific code to be played next and output as block 324. In some examples, artists can provide information for each code, such as its weight and how often that particular code should be used.

[0087] Next, in block 326, chord variations are selected, at least in part, based on a harmony complexity parameter that serves as an input. The harmony complexity parameter 328 can be adjusted or selected by the artist, or it can be determined automatically. Generally, a higher harmony complexity parameter may be associated with a higher energy audio output, and a lower harmony complexity parameter may be associated with a lower energy audio output. In some cases, the harmony complexity parameter may include inputs such as chord inversion, vocalization, and harmony density.

[0088] In block 330, the process obtains the root of the code, and in block 332, selects the bass segments to be played from the available bass segments 334. These bass segments then undergo bus processing 336, where they can be equalized, filtered, timed, and otherwise processed.

[0089] Returning to the chord variation in block 326, the process continues separately to block 338, where a selected harmony is played from the available harmony segments 340. This harmony segment then undergoes bus processing 342. Harmony segment bus processing 342 allows for processing such as equalization, filtering, and timing, similar to bass bus processing.

[0090] Returning to the selected code 324, the process continues to filter melody notes individually in block 344 using the input of the melody constraint 346. The output in block 348 is the melody notes available for playback. The melody constraint 346 can be provided by the artist and may, for example, specify which notes to play or not, limit the melody range, or provide other such constraints that may depend on a particular selected code 324.

[0091] In block 350, the process determines which melody note to play (from among the available melody notes 348). This decision can be made automatically based on model values, artist-provided inputs, randomization effects, or any other suitable inputs. In the illustrated example, one input is from trigger melody note block 352, which is based on a melody density parameter 354. The artist can provide a melody density parameter 354, which partially determines how complex and / or high-energy the audio output is. Based on that parameter, melody notes may be triggered more or less frequently and at specific times to determine which melody note to play using block 352, which is input to block 350. In various examples, the output of block 350 may be provided as an input to block 350 in the form of a feedback loop, such that the next melody note selected in block 350 depends at least partially on the last melody note selected in block 350. Next, in block 356, a melody segment is selected from among the available melody segments 358, and then the melody segment undergoes bus processing 360.

[0092] Returning to the beginning of block 302, the process proceeds separately to block 362 to play non-musical content. This may be, for example, natural sounds, spoken language, or other such non-musical content. Various non-musical segments 364 can be stored and played back. These non-musical content segments may also undergo bus processing in block 366.

[0093] The outputs of these various paths (e.g., selected bass segments(s), harmony segments(s), melody segments(s), and / or non-musical segments(s)) can each undergo separate bus processing before being combined in block 368 via mixing and mastering. Here, combined levels can be set, various filters can be applied, relative timings can be established, and any other appropriate processing steps can be performed before the generated audio content is output in block 370. In various examples, some of the paths may be omitted entirely. For example, the generated media module may omit the option to play non-musical content together with the generated music content. The process 300 shown in Figure 3 is illustrative, and as those skilled in the art will see, appropriate modifications can be made to the process 300 shown herein, and there are numerous other appropriate alternative processes that can be used to generate the generated media content.

[0094] Figure 4 shows an exemplary architecture for storing and retrieving generated media content. In this example, the generated media content includes various individual tracks (each with multiple variations related to energy levels or other parameters) that can be selected and played in various orders and groupings depending on specific input parameters.

[0095] As shown in the diagram, the generated media content 404 can be stored as one or more audio files associated with the global generated media content metadata 402. Such metadata may include, for example, the global tempo (e.g., number of beats per minute), the global trigger frequency (e.g., how often to check for changes in input parameters), and / or the global crossfade duration (e.g., the time to fade between different selected energies).

[0096] The generated media content 404 contains several different tracks 406, 408, and 410. During operation, these tracks can be selected and played in various arrangements (e.g., randomized grouping by some overlay, or playback according to a predetermined sequence). In some examples, the generated media content 404, including tracks 406, 408, and 410, can be stored locally via one or more playback devices, while one or more remote computing devices can periodically transmit updated versions of the tracks, generated media content, and / or global generated media content metadata. In some examples, the remote computing devices can be periodically polled or queried by the playback devices, and in response to the query or polling, the remote computing devices can supply updates to the generated media modules stored in the local playback devices.

[0097] For each track, there may be a corresponding subset of that track corresponding to a different energy level. For example, the first energy level (EL) of track 1 in 412, the second energy level of track 1 in 414, and the n energy level of track 1 in 416. Each of these may include both metadata corresponding to a particular energy level (e.g., metadata 418, 420, 422) and a specific media file (e.g., media files 424, 426, 428). In some examples, each track may contain multiple media files (e.g., media file 424) arranged in a particular way, and their arrangement and combination may be desired by the corresponding metadata (e.g., metadata 418). The media files can be in any suitable format that can be played back, for example, via a playback device and / or streamed to a playback device for playback. In some examples, one or more of media files 424, 426, 428 may be the output of the generative model shown in Figure 3. Metadata may include, for example, tempo (if different from global tempo), trigger frequency (if different from global trigger frequency), sequence information (e.g., whether specific files are played sequentially, randomly, or percentage-weighted), crossfade duration (if different from global crossfade), spatial information (e.g., for rendering audio content in space using multiple transducers), polyny information (e.g., to allow multiple audio files to be played simultaneously in this segment), and / or levels (e.g., level adjustment in dB units or random within a given range).

[0098] During operation, a target energy level can be determined using one or more input parameters (e.g., the number of people in the room, the time of day). This determination can be made using a playback device and / or one or more remote computing devices. Based on this determination, specific media files corresponding to the determined energy level can be selected. The generating media module can then arrange and play these selected tracks according to the generated content model. This may include playing the selected tracks in a specific predetermined order, in a random or pseudo-random order, or any other appropriate method. The tracks can be played in a way that overlaps at least partially, in some examples. It may be useful to vary the amount of overlap between tracks so that a casual listener does not hear a repeating loop of audio content, but instead perceives the generated audio as an infinite stream of audio without repetition.

[0099] The example shown in Figure 4 utilizes energy levels as a parameter to distinguish between different generated audio content, but in various examples, specific variations or permutations of generated audio content may change along other dimensions (e.g., genre, time of day, associated user task, etc.).

[0100] d. Exemplary sensor data sources and other input parameters As described above, the generation media module 214 can generate generation media based at least in part on input parameters that can include sensor data (for example, if received from the sensor data source 218) and / or other appropriate input parameters. With respect to sensor input parameters, the sensor data source 218 can include data from any appropriate sensor, wherever it is located relative to the generation media group, whether it is the value measured thereby. Examples of appropriate sensor data include physiological sensor data, such as data obtained from biometric sensors, wearable sensors, etc. Such data may include physiological parameters such as heart rate, respiratory rate, blood pressure, electroencephalogram, activity level, movement, and body temperature.

[0101] Suitable sensors include wearable sensors configured to be worn or carried by a user, such as headsets, watches, mobile devices, brain-machine interfaces (e.g., Neuralink), headphones, microphones, or other similar devices. In some examples, the sensor may be a non-wearable sensor or may be fixed to a fixed structure. The sensor may provide sensor data that may include, for example, data corresponding to brain activity, voice, position, movement, heart rate, pulse, body temperature, and / or sweating. In some examples, the sensor may correspond to multiple sensors. For example, as described elsewhere in this specification, the sensor may correspond to a first sensor worn by a first user, a second sensor worn by a second user, and a third sensor not worn by a user (e.g., fixed to a stationary object or structure). In such examples, the sensor data may correspond to multiple signals received from each of the first, second, and third sensors.

[0102] The sensor can be configured to acquire or generate information that generally corresponds to the user's mood or emotional state. In one example, the sensor is a wearable brain-sensing headband, which is one of many examples of the sensors described herein. Such a headband could include, for example, an electroencephalogram (EEG) headband having multiple sensors on it. In some examples, the headband could correspond to one of the Muse® headbands (InteraXon; Toronto, Canada). The sensors can be positioned at various locations around the inner surface of the headband to correspond, for example, to different anatomical structures of the user's brain (e.g., the frontal, parietal, temporal, and sphenoid bones). Thus, each sensor can receive different data from the user. Each sensor can correspond to an individual channel that can be streamed from the headband to system devices 210 and / or 250. Such sensor data can be used to detect the user's mood, for example, by classifying the frequencies and intensities of various brain waves, or by performing other analyses. Further details regarding the use of a brain-sensing headband for generated audio content can be found in co-owned U.S. Patent Application No. 62 / 706,544, filed on 24 August 2020, entitled MOOD DETECTIONAND / OR INFLUENCE VIA AUDIO PLAYBACK DEVICES, which is incorporated herein by reference in its entirety.

[0103] In some examples, the sensor data source 218 includes data acquired from network devices (e.g., Internet of Things (IoT) sensors such as networked lights, cameras, temperature sensors, thermostats, presence detectors, and microphones). In addition to or instead of this, the sensor data source 218 may include environmental sensors (e.g., measuring or displaying weather, temperature, time / day / week / month, etc.).

[0104] In some examples, the generating media module 214 may utilize inputs in the form of playback device capabilities (e.g., number and type of transducers, output power, other system architecture) and device location (e.g., location relative to one or more users, or any location relative to other playback devices). Further examples of generating and modifying generated audio as a result of user and device location are described in detail in U.S. Patent Application No. 62 / 956,771, filed 3 January 2020, titled “GENERATIVE MUSIC BASED ON USER LOCATION,” owned by the same owner, which is incorporated herein by reference in its entirety. Additional inputs may include device states of one or more devices in a group, such as thermal conditions (e.g., if a particular device is at risk of overheating, the generated content may be modified to lower the temperature), battery levels (e.g., bass output may be reduced in portable playback devices with low battery levels), and coupling states (e.g., whether a particular playback device is configured as part of a stereo pair, coupled with a subwoofer, or configured as part of a home theater system). Any other suitable device characteristics or states can similarly be used as input for generating the generated media content.

[0105] Another exemplary input parameter could include the presence of a user. For example, when a new user enters a space where generated audio is being played, the system could detect the user's presence (e.g., via proximity sensors, beacons, etc.) and modify the generated audio in response. This modification could be based on the number of users (e.g., ambient medicinal audio for one user, relaxing music for two to four users, and party or dance music for more than four users). The modification could also be based on the identification information of the present users (e.g., a user profile based on user characteristics, listening history, or other such identifiers).

[0106] For example, a user may wear a biometric device that can measure various biometric parameters, such as the user's heart rate or blood pressure, and report those parameters to devices 210 and / or 250. The generating media module 214 of these devices 210 and / or 250 can use these parameters to further adapt the generated audio, for example, by increasing the tempo of the music in response to the detection of a high heart rate (so as to indicate that the user is engaged in high-intensity physical activity), or by decreasing the tempo of the music in response to the detection of high blood pressure (so as to indicate that the user is stressed and could benefit from calming music).

[0107] In yet another example, one or more microphones in a playback device (e.g., microphone 115 in Figure 1F) can detect the user's voice. The captured voice data can then be processed to determine, for example, the user's mood, age, or gender, to identify a specific user from among multiple users in the household, or to identify any other such input parameters. Other examples are similarly possible.

[0108] e. Examples of coordination among group members Figure 5 is a functional block diagram illustrating data exchange in a system for playback of generated media content. For illustrative purposes, the system 500 shown in Figure 5 includes interaction between a coordinator device 210 and a member device 250b. However, the interactions and processes described herein can be applied to interactions including a plurality of additional coordinator devices 210 and / or member devices 250. As shown in Figure 5, the coordinator device 210 includes a generated media module 214a that receives inputs including input parameters 502 (e.g., sensor data, media content, model parameters of the generated media module 214a, or other such inputs) and clock and / or timing data 504. In various examples, the clock and / or timing data 504 may include synchronization signals to synchronize playback and / or synchronize generated media being generated by various devices in the group. In some examples, the clock and / or timing data 504 may be provided by an internal clock, a processor, or other such component housed in the coordinator device 210 itself. In some cases, the clock and / or timing data 504 can be received from a remote computing device via a network interface.

[0109] Based on these inputs, the generating media module 214a can output generated media content 404a. Optionally, the output generated media content 404a itself can function as an input to the generating media module 214a in the form of a feedback loop. For example, the generating media module 214a can generate subsequent content (e.g., audio frames) using a model or algorithm that at least partially depends on previously generated content.

[0110] In the illustrated example, member device 250b similarly includes a generating media module 214b which may be substantially the same as the generating media module 214a of coordinator device 210, or may differ in one or more embodiments. The generating media module 214b can also receive input parameters 502 and clock and / or timing data 504. These inputs can be received from coordinator device 210, from other member devices, from other devices on the local network (e.g., a locally networked smart thermostat supplying temperature data), and / or from one or more remote computing devices (e.g., a cloud server providing clock and / or timing data 504, or weather data, or any other such inputs). Based on these inputs, the generating media module 214b can output generated media content 404b. This generated generated media content 404b can optionally be fed back to the generating media module 214b as part of a feedback loop. In some cases, the generated media content 404b may include or consist of the generated media content 404a transmitted to the member device 250b via the network (generated via the coordinator device 210). In other cases, the generated media content 404b may be generated separately and independently of the generated media content 404a generated via the coordinator device 210.

[0111] The generated media contents 404a and 404b can then be played back via devices 210 and 250b themselves and / or by other devices in the group. In various examples, the generated media contents 404a and 404b can be configured to be played back simultaneously and / or synchronously. In some cases, the generated media contents 404a and 404b may be substantially identical or similar to each other, and each generated media module 214 utilizes the same or similar algorithms and the same or similar inputs. In other examples, the generated media contents 404a and 404b may still be configured for synchronous or simultaneous playback, but may be different from each other.

[0112] f. Exemplary generation media using a distributed architecture As mentioned earlier, media content generation can be computationally intensive, and in some cases, it may be impractical to perform it entirely on a local playback device alone. In some examples, a local playback device's generating media module can request generated media content from generating media modules stored on one or more remote computing devices (e.g., cloud servers). The request may include or be based on specific input parameters (e.g., sensor data, user input, contextual information). In response to the request, the remote generating media module can stream specific generated media content to the local device for playback. The specific generated media content provided to the local playback device may change over time depending on specific input parameters, the configuration of the generating media module, or other such parameters. In addition to or instead of this, the playback device may store individual tracks for playback (e.g., different variations of tracks are associated with different energy levels, as shown in Figure 4). The remote computing device can then periodically provide the local playback device with new files for updated tracks for playback, or it can provide updates to the generating media module that decides when and how to play specific files stored locally on the playback device.

[0113] In this way, the tasks required to generate and play the generated audio are distributed between one or more remote computing devices and one or more local playback devices. Overall efficiency can be improved by performing at least some of the computationally intensive tasks associated with generating new media content on the remote computing devices, and optionally by reducing the need for real-time computation. By generating a discrete number of alternative tracks or track variations via the remote computing devices according to a specific media content model before playback, the local playback devices can request and receive specific variations based on real-time or near-real-time input parameters (e.g., sensor data). For example, the remote computing devices can generate different versions of media content, and the playback devices can request a specific version in real-time based on input parameters. As a result, playback of appropriate generated media content based on real-time or near-real-time input parameters (e.g., sensor data) is performed without requiring de novo generation of such media content to be performed in real time.

[0114] Figure 6 is a schematic diagram of an exemplary distributed generated media playback system 600. As shown, an artist 602 can supply multiple media segments 604 and one or more generated content models 606 to a stored generated media module 214 via one or more remote computing devices. The media segments may correspond to, for example, specific audio segments or seeds (e.g., individual notes or chords, short tracks of n bars, non-musical content, etc.). In some examples, the generated content models 606 may also be supplied by the artist 602. This may include providing the entire model, or the artist 602 may provide input to the model 606 by, for example, changing or adjusting specific aspects (e.g., tempo, melody constraints, harmony complexity parameters, chord change density parameters, etc.).

[0115] The generating media module 214 can receive both a media segment 604 and one or more input parameters 502 (as described elsewhere in this specification). Based on these inputs, the generating media module 214 can output generated media. As shown in Figure 6, the artist 602 can audition the generating media module 214 by optionally receiving exemplary outputs based on inputs provided by the artist 602 (e.g., media segment 604 and / or generated content model 606). In some cases, the audition can play back variations of generated media content to the artist 602 depending on various different input parameters (e.g., one version corresponding to high energy levels intended to produce an exciting or uplifting effect, and another version corresponding to low energy levels intended to produce a calming effect). Based on the output from this audition step, the artist 602 can dynamically update the settings of media segment 604 and / or generated content model 606 until the desired output is achieved.

[0116] In the illustrated example, there may be iterations in block 608 every n hours (or minutes, days, etc.) in which the generating media module 214 can generate multiple different versions of the generated media content. In the illustrated example, there are three versions: version A in block 610, version B in block 612, and version C in block 614. These outputs are stored as generated media content 616 (e.g., via a remote computing device). A particular version (version C as block 618 in this example) can be sent to a local playback device 250 for playback (e.g., streamed). In some examples, a particular version may correspond to tracks 406, 408, and 410 shown in Figure 4.

[0117] While three versions are shown here as examples, in reality, many more versions of generated media content produced via remote computing devices may exist. These versions can vary along several different dimensions, such as being suitable for different energy levels, different intended tasks or activities (e.g., study vs. dance), different times of day, or any other appropriate variations.

[0118] In the illustrated example, the playback device 250 can periodically request a specific version of the generated media content from a remote computing device. Such requests may be based, for example, on user input (e.g., user selection via a controller device), sensor data (e.g., the number of people present in a room, background noise level, etc.), or other appropriate input parameters. As illustrated, the input parameter 502 may be optionally provided (or detected) to the playback device 250. In addition to or instead of this, the input parameter 502 may be provided (or detected) to the remote computing device 106. In some examples, the playback device 250 transmits the input parameter to the remote computing device 106, and the remote computing device provides the playback device 250 with an appropriate version without the playback device 250 specifically requesting a particular version. g. Examples of methods for generating digital content based on blockchain data

[0119] As mentioned above, a system for generating and playing generated media content may be able to interact with blockchain data (or data stored via other distributed ledger technologies). For example, as shown in Figure 6, the blockchain layer 620 can be used to provide data as input to other components of the generated media playback system 600, such as the playback device 250, input parameters 502, media segments 604, generated content models 606, and / or generated media modules 214. In various embodiments, the blockchain layer 620 may store data that can be used as one or more input parameters 502, data that can be included or used to obtain or influence a particular media segment 604, data that can be included or used to obtain or influence a particular generated content model 606, and / or data that can be included or used to obtain or influence a particular generated media module 214. Furthermore, some or all of these components can communicate with the blockchain layer 620 to write data to the blockchain, record transactions, and even interact with the blockchain layer 620. For example, the playback device 250 may record a transaction that reflects the playback of a particular track on the blockchain layer 620. Data stored via the blockchain layer 620 may include specific input parameters 502, or appropriate input parameters 502 may be generated using data stored via the blockchain layer 620. Similarly, specific media segments 604, generated content models 606, generated media modules 214, input parameters 502, or other appropriate data may be written to the blockchain layer 620 to create an immutable record of such content, transactions, or other data. Further details regarding the use of blockchain technology (or other appropriate distributed ledger technology) in the creation and playback of generated media content are described in more detail below.

[0120] Figure 7 is a schematic diagram of another example of a distributed generated media playback system 700. As shown, the system 700 includes, or communicates with, a blockchain layer 620 to obtain, generate, or store, for example, input parameters 502, a generated content model 606, or other data or parameters used in the creation and playback of generated media content.

[0121] Examples of such blockchain layer 620 include public distributed ledgers such as Ethereum, Bitcoin, Solana, Avalanche, and Polygon. While blockchain layer 620 is illustrated, in various implementations, any suitable distributed ledger technology can be used, including private or semi-private blockchains, as well as non-blockchain implementations such as directed acyclic graphs (DAGs) (e.g., Nano, IOTA, etc.). In various examples, participants using blockchain layer 620 can trade with each other in a peer-to-peer manner, and the operation of blockchain layer 620 can be decentralized so that no single central entity controls the operation of the network. Such distributed ledgers can be used to track the creation, exchange, and redemption of certain real-world assets, such as currencies. This approach allows for robust auditing of asset transactions due to the practical immutability of the data stored on the blockchain. Currency is just one of many assets that are desirable to track on a distributed ledger. Other types of assets may differ from currency in one or more actions that govern the creation, exchange, and / or redemption of the asset. Furthermore, different blockchain architectures may differ in terms of policies, protocols, and even the tools used to program asset behavior.

[0122] Generally, a distributed ledger where each unit of an asset is represented by some form of digital token can be programmed to impose a set of behaviors appropriate to the asset that the token represents. For example, "fungible" behavior allows an asset to be exchanged for other assets of the same class. A unit of a certain denomination of currency (e.g., $1) is fungible because it has the same value as other units of the same denomination of currency. In contrast, a property title is "non-fungible" because its value depends on the size, location, and other aspects of the specified property. For each asset represented as a token, the appropriate fungible or non-fungible behavior is programmed into the class of tokens in the virtual ledger that tracks the asset.

[0123] In some embodiments, tokens traded over blockchain layer 620 are nonfungible. Such nonfungible tokens (NFTs) are unique and cannot be exchanged for other tokens. NFTs may consist of and / or be associated with unique digital artwork and / or music, domain names, digital collectibles (such as CryptoKitties or memes), event tickets, parts of virtual worlds, digital objects used in games, avatars or characters, items with utility (such as providing voting rights or governing rights or other specific functions). In various examples, the NFT itself may contain associated data (for example, the raw audio data of a music NFT may be stored on-chain), or the NFT may contain a pointer (such as a URL or URI) that directs to data stored elsewhere (for example, audio data stored on a server managed by the issuer of the music NFT).

[0124] Such tokens, whether fungible or nonfungible, can be stored by users via digital wallets. A digital wallet is a device, physical medium, program, or service that can store the public and / or private keys of blockchain transactions. In some examples, a digital wallet can store multiple public / private key pairs for various different blockchains, allowing users to store assets related to different blockchains in a single wallet. Examples include MetaMask, Phantom, Coinbase Wallet, and Ledger Nano. In operation, users can sign blockchain transactions via a wallet using the appropriate private key (or by allowing the wallet to sign transactions with the private key). The transaction is then verified if the transaction signature is valid and added to the corresponding block on the blockchain. In some examples, a wallet identifier may be used as an input parameter to a generative model, independently of the tokens held in a particular wallet.

[0125] In various implementations, the blockchain layer 620 can be configured to automatically execute transactions under one or more conditions. Such self-executing transactions may be called “smart contracts.” Smart contracts are computer code stored on the blockchain and are configured to execute only under specific circumstances or in specific ways. For example, a smart contract can be configured to execute a specific transaction at a certain time if a certain threshold is exceeded based on one or more other transactions or other appropriate criteria. In some examples, the generated media module 214 and / or generated content model 606 can be implemented in the form of a smart contract, so that by interacting with the smart contract via the blockchain layer 620, the smart contract can output generated media content, a generated content model, or data or instructions that can be used to output such generated media content or to generate a generated content model.

[0126] One organizational structure unique to blockchain is the Decentralized Autonomous Organization (DAO). A DAO is typically a community-driven entity without a central authority. Such a DAO becomes fully autonomous and transparent because smart contracts provide the underlying rules and execute agreed-upon decisions. Community voting can be conducted by token holders using on-chain transactions. Based on specific voting results, smart contracts can execute specific transactions or other code to implement decisions made by DAO members. Generally, DAOs issue tokens to users in exchange for currency investments or donations, or free of charge (e.g., through "airdrops"). Token holders typically hold a certain amount of voting rights, which may be proportional to the number of tokens they hold. In some cases, token holders also receive monetary returns, such as a share of transaction fees collected by the DAO.

[0127] In the exemplary system 700 shown in Figure 7, the generating media module 214 can receive a number of different inputs and output one or more generated content versions 610 in response. These content versions 610 are stored in the generating media content store 616, from which a specific selected generated content version 618 can be selected and played back via the playback device 250 or other output device (for example, the light component of the generated media content version 618 can cause the lighting device 702 to output light, in which case specific hue, color temperature, brightness, on / off or other patterns etc. will follow the generated media content version 618). The inputs to the generating media module 214 include a generated content model 606 and input parameters 502, similar to the approach described above with respect to Figure 6. As previously mentioned, in one embodiment, the generated content model 606 may be stored via the blockchain layer 620 or obtained from data stored via the blockchain layer 620. Similarly, one or more input parameters 502 may include or be based on data stored via the blockchain layer 620. For example, blockchain data can be used in the generating media module 214 to generate appropriate outputs, such as "sonicing" the data stream (for example, a real-time feed such as cryptocurrency price values ​​can be converted into a corresponding sound output). In some examples, the blockchain data may include data provided by one or more "oracles," which are typical third-party services that provide external information (e.g., price feeds, weather data, election results, etc.) to smart contracts. In some examples, the generating media module 606 and / or the generated media content 616 may be stored locally via the playback device 250, in which case the input parameters 502 are streamed to the playback device 250 and used to generate a new version 610 of the generated media content 616.

[0128] Additionally, or alternatively, the generating media module 214 can receive the output of one or more smart contracts 706 as input. In some cases, the generating media module 214 can take the form of a smart contract itself, in which case the program code is stored on the blockchain and executed automatically under certain conditions (for example, a user 708 interacts with the smart contract 706, and in response, a specific generated media content is sent to a specified destination). In some examples, the generating media module 214 can run locally (rather than as a smart contract on the blockchain) or via a remote server, but can communicate with the smart contract 706 to receive input parameters 502 from the smart contract 706 or provide appropriate output to the smart contract 706. For example, a specific generated media content output by the generating media module 214 can be used to generate one or more NFTs via the smart contract 706. In some embodiments, each specific version of the generated content produced by the generating media module 214 can take a corresponding NFT, thereby ensuring that each NFT produced by the smart contract 706 based on input from the generating media module 214 is unique. This is illustrated in Figure 7, where multiple discrete NFTs 710a to f, represented by NFT1 to NFTn, are generated via the smart contract 706. These NFTs 710 can further be provided as input to the generating media module 214. For example, the generating media module 214 can dynamically generate different content based at least partially on the specific NFTs 710 it interacts with. Additionally, or instead, a user 708 can access a specific generating media module 214 only if the user holds the appropriate NFTs 710 in their digital wallet.

[0129] In some examples, one or more NFTs 710 are existing NFTs owned by third-party individuals or entities (e.g., individuals or entities unrelated to user 708) and are temporarily accessible by the generating media module 214. In certain examples, one or more NFTs 710 do not consist of audio data, but instead consist of other data types (e.g., video, images, or other data), and these data are either "sound-wave-ified" or converted into a form that can generate media content by the generating media module 214 (or another appropriate component).

[0130] In the illustrated example, smart contract 706 can also output NFT712, represented by NFT0 held by user 708. Furthermore, this NFT712 can interact with or be generated through artist DAO714, and this artist DAO714 can communicate with one or more smart contracts 706. As mentioned earlier, DAOs are generally community-driven organizations where members hold tokens (e.g., NFT712), and the tokens designate membership, provide voting rights and other governing rights, and additionally grant token holders economic benefits such as earnings from future DAO revenue. In some examples, artist DAO714 can distribute a portion of royalties (music usage fees) to holders of appropriate NFTs or other tokens (for example, user 708 can receive economic benefits from artist DAO714 based at least partially on their ownership of user NFT712).

[0131] Optionally, data corresponding to the NFT712 can be stored or embedded via a physical medium. For example, as shown in Figure 7, data corresponding to the NFT712 can be embedded in the record 716 (e.g., via a unique QR code®, a code embedded in the grooves of the record 716), or the data can be stored using any other suitable technology via another physical medium (e.g., NFC or other RF tags). Although the record 716 is illustrated, in various embodiments, the physical substrate can take various forms, such as a playback device, a physical card or ticket, or a poster.

[0132] By using blockchain layer 620, smart contracts 706, DAO 714, and / or NFTs 710 and 712, several advantages can be provided to the generated media playback system 700. For example, by associating a specific NFT with generated media content (e.g., a soundscape), a user 708 can obtain a personalized history, which can then be transferred via the distributed peer-to-peer transaction mechanism of blockchain layer 620. However, a problem with this approach is that the data contained in NFTs is generally static, which is in conflict with the dynamic data of generated soundscapes and other generated media content. Another problem associated with NFTs is broken links, where locators within an NFT cease to reference the artwork associated with the NFT, and the data within the NFT becomes outdated. One way to mitigate the possibility of broken links is to store the generated media content or the generated media engine on blockchain layer 620. Furthermore, the artwork itself may be embedded in the NFT (e.g., artwork data is stored on-chain rather than on a separate server).

[0133] In some cases, the NFT 710 may contain a specific seed used by the generating media module 214. Examples of such seeds include the media segment 604 in Figure 6, the track, energy level, or metadata in Figure 4, or any of the various components shown in Figure 3. Optionally, the characteristics of a particular NFT (at least with respect to its use by the generating media module 14) may depend on its transaction history. For example, a specific seed associated with an NFT may change dynamically depending on when it was last traded and how many times it has been traded. In addition, or alternatively, different combinations of NFT 710 connected to the generating media module 214 may result in different generated content versions 610, and a particular generated media output depends on which of the NFT 710 is used as input.

[0134] In some examples, additional data related to the generated media playback system 700 can be stored via the blockchain layer 620. For example, a user 708's listening history can be stored in the blockchain layer 620 to provide an immutable record of the listening history. This could include viewing history of generated media content or non-generated content (such as standard pre-recorded audio tracks or other content). In some cases, a "follower" of a particular user 708 can subscribe to that user's listening history by accessing the data stored via the blockchain layer 620. Because blockchains are generally permissionless and transparent, followers can freely access the listening history (or other content data associated with a particular network address). In yet another example, a follower can subscribe to a particular user 708's generated media content. This would allow user 708's generated media content to be dynamically created based on various inputs, but this same media content can be enjoyed by other followers.

[0135] For example, consider an artist who wants to create a specific soundscape using a generated media module 214. The artist's fans can listen to the soundscape in real time or near real time via blockchain data. In some cases, followers can use their own local generated media module 214 (or one running on another device) to generate corresponding generated media content using inputs, pointers, or other data. In at least some examples, such a local generated media module 214 can also utilize additional local inputs (e.g., specific playback device characteristics, local sensor data, etc.). As a result, the artist's generated media content is merged with that generated by the user's own local generated media module 214, producing generated media content that reflects the artist's intentions but is slightly modified. Optionally, the artist's followers may be required to hold a specific NFT or other token in order to access the artist's generated media content. In yet another example, certain playlists or radio stations may only be accessible to users who hold a specific NFT or token.

[0136] h. Exemplary methods for generating and playing back generated audio Figures 8 to 13 are flowcharts illustrating exemplary methods for playing generated audio content through multiple separate playback devices. Methods 800, 900, 1000, 1100, 1200, and 1300 can be implemented using any of the devices or systems described herein, or any other devices or systems currently known or to be developed in the future.

[0137] Various examples of Methods 800, 900, 1000, 1100, 1200, and 1300 include one or more operations, functions, or actions represented by blocks. Although the blocks are shown in a sequential order, these blocks may also be executed in parallel and / or in an order different from the order disclosed and described herein. Furthermore, the various blocks may be combined into fewer blocks, divided into additional blocks, and / or removed, based on the desired implementation form.

[0138] Furthermore, with respect to Methods 800, 900, 1000, 1100, 1200, 1300, and other processes and methods disclosed herein, flowcharts illustrate the functions and operations of several example possible implementations. In this regard, each block may correspond to a module, segment, or part of program code containing one or more instructions that can be executed by one or more processors to perform a particular logical function or step in a process. The program code may be stored in any type of computer-readable medium, such as a storage device including, for example, a disk or hard drive. The computer-readable medium may include non-temporary computer-readable media such as, for example, register memory, processor cache, and tangible non-temporary computer-readable media that store short-term data, such as random access memory (RAM). The computer-readable medium may also include non-temporary media such as, for example, read-only memory (ROM), optical or magnetic disks, and compact disk read-only memory (CD-ROM), which are secondary or persistent long-term storage devices. The computer-readable medium may also be any other volatile or non-volatile storage system. A computer-readable medium may be considered, for example, a computer-readable storage medium or a tangible storage device. Furthermore, with respect to the methods disclosed herein and other processes and methods, each block in Figures 8 to 13 may correspond to a circuit wired to perform a specific logical function within the process.

[0139] Referring to Figure 8, method 800 begins in block 802, which includes receiving a command to play generated media content via a group or combined zone of playback devices. Such a command may be received, for example, via a control device 130 or other appropriate user input.

[0140] In block 804, method 800 includes a group coordinator device providing timing information to a generating group member device. The timing information may include context timing data (e.g., time data associated with a sensor input or other user input), generated media playback timing data (e.g., timestamps and synchronization data to facilitate synchronous playback of generated media), and / or media content stream timing data based on a common clock.

[0141] In block 806, the method optionally includes determining a generated media content model to be used to generate the generated media. Such a model can be implemented, for example, in the media content module 214 described with respect to Figures 2 to 6. In some examples, each member device may utilize the same or substantially the same generated media content model, while in other cases, some or all of the member devices may utilize different generated media content models. For example, a first generated media content model may generate rhythmic beats, and a second generated media content model may generate ambient natural sounds. When played simultaneously, the generated audio produced by these different generated media content models can produce a pleasant listening experience for the user. In some examples, the selection of a particular generated media content model itself may be based on one or more input parameters such as device capabilities, device location, the number of users present, and user sensor data.

[0142] In block 808, method 800 includes the coordinator device and member devices receiving context and / or other input data. For example, the input data may include sensor data, user input, context data, or any other relevant data that can be used as input for a generated media content model.

[0143] Method 800 continues in block 810 for the coordinator device and member devices to synchronously generate and play generated media content.

[0144] Figure 9 shows another method 900 for playing generated audio content through multiple playback devices. Method 900 begins in block 902 with a group coordinator device receiving one or more input parameters. As previously mentioned, the input parameters may include sensor data, user input, context data, or any other inputs that may be used by the generating media module to generate generated audio for playback.

[0145] In block 904, the coordinator device transmits input parameters to one or more discrete playback devices having a generated media module. For example, the coordinator device may acquire sensor data and other input parameters and transmit them to multiple individual playback devices in an environment, or to multiple individual playback devices distributed across multiple environments. In some examples, these input parameters may include features of the generated content model itself, for example, providing instructions to update a generated media module stored locally by one or more of the separate playback devices.

[0146] In block 906, the method includes transmitting timing data from a coordinator device to a playback device. The timing data may include, for example, clock data or other synchronization signals configured to facilitate the coordination of the generation of generated media content and the synchronous playback of that generated media content via a separate playback device.

[0147] Method 900 continues in block 908 to simultaneously play generated media content through playback devices based at least in part on input parameters. As previously stated, the various playback devices may play the same generated audio, or each may play separate generated audio that, when played synchronously, produces a psychoacoustic effect desired by the present user.

[0148] In the example in Figure 9, the generated media content can be generated locally by discrete playback devices, each of which generates and plays its own generated audio content in parallel with the others. In another method 1000 shown in Figure 10, the generated media content is generated by a coordinator device, which then transmits the generated media content along with timing data to separate playback devices for synchronized playback.

[0149] In block 1002, method 1000 includes receiving one or more input parameters in a group coordinator device. Examples of input parameters are described elsewhere in this specification and include sensor data, user input, context data, or any other inputs that may be used by a generating media module to generate generated audio for playback.

[0150] In block 1004, the coordinator device generates first and second generated media streams based at least partially on input parameters, and in block 1006, the first and second media streams are sent to the first and second separate playback devices, respectively. For example, the coordinator device can generate two streams that form different channels of generated audio, for example, having a left channel played by the first playback device and a corresponding right channel played by the second playback device. In addition to or instead of this, the two streams can be separate audio tracks that can nevertheless be played synchronously, such as a rhythmic beat in one stream and ambient nature sounds in the other stream. Several other variations are possible. This example describes two streams for two playback devices, but in various other examples, there may be one or more streams that can be provided to any number of playback devices for synchronous playback. In at least some examples, one or more of the playback devices may be located in different environments far apart from each other (e.g., different homes, different cities, etc.).

[0151] In block 1008, the first playback device plays the first generated media stream, and the second playback device plays the second generated media stream. In some examples, this simultaneous playback can be facilitated by using timing data received from a coordinator device.

[0152] Figure 11 shows another exemplary method 1100 for generating and playing generated media content. As previously mentioned, it may be beneficial to use one or more remote computing devices (e.g., cloud-based servers) to perform at least some of the processing required to generate the generated media content, thereby reducing the computational demands imposed on the local playback device and / or performing operations that would be impossible using the components of the local playback device. Method 1100 begins in block 1102 with the playback device receiving one or more input parameters. As previously mentioned, the input parameters may include sensor data, user input, context data, or any other inputs that may be used by the generated media module to generate generated audio for playback.

[0153] In block 1104, method 1100 includes accessing a library containing multiple existing media segments. For example, multiple individual media segments (e.g., audio tracks) can be stored on a playback device and arranged and / or mixed for playback according to a generated content model. In addition to or instead of this, the library can be stored on one or more remote computing devices, and the individual media segments are sent from the remote computing devices to the playback devices for playback.

[0154] Method 1100, in block 1106, proceeds to generate media content, at least in part, based on input parameters, by arranging a selection of existing media segments from a library for playback according to a generated media content model. As described elsewhere in this specification, a generated media content model can take one or more input parameters as input. Based on the input, a specific generated media content can be output using the generated media content model. In the example, the generated media content may include an arrangement of existing media segments, for example, with or without overlap between specific media segments, and / or with additional processing or mixing steps performed to produce the desired output.

[0155] In block 1108, the playback device plays the generated media content. In various examples, this playback can be performed simultaneously with and / or in sync with additional playback devices.

[0156] Figure 12 shows Method 1200, another example for generating and playing generated media content. As mentioned above, it is beneficial to incorporate or rely on blockchain data to produce generated media content. Method 1200 begins in block 1202 by accessing blockchain data stored via a distributed ledger through a playback device. The distributed ledger can be a public blockchain such as Ethereum, Bitcoin, or Solana, or optionally a private or semi-private blockchain, or a non-blockchain ledger. The blockchain data can include one or more existing media segments or other seeds used to generate the generated media content. In some examples, such data is stored directly on the blockchain itself, while in other examples, the blockchain can store pointers (e.g., URLs or URIs) that indicate where the media segments or other seed data are stored. Optionally, the blockchain data can take the form of one or more non-fungible tokens (NFTs).

[0157] Method 1200 generates media content in block 1204 via a playback device, at least partially based on blockchain data. In some examples, this generation involves accessing a library of existing media segments stored in the playback device or other suitable storage location (e.g., a remote server, another device on the local network). In some cases, these media segments are retrieved from the blockchain or other remote locations and stored via the playback device. The playback device can then arrange a selection of existing media segments from the library for playback according to a generated media content model. This selection can be at least partially based on blockchain data. For example, as described elsewhere in this specification, a generated media content model can communicate with a smart contract or a decentralized autonomous organization (DAO) in a way that influences specific generated media content output by the model. In some cases, specific NFTs or other tokens may influence the output of a generated media content model. Such NFTs or other tokens may, in some cases, be used in combination so that a particular combination of NFTs or other tokens generates a unique output via the generated media content model. In at least some cases, two or more blockchains can be used simultaneously in this manner (for example, the generated media content model can vary its output based on a user who holds both a first NFT on the Solana network and a second NFT on the Ethereum network). In block 1206, method 1200 includes playing the generated media content via a playback device.

[0158] Figure 13 shows another exemplary method 1300 for generating and playing generated media content. As mentioned above, smart contracts and other self-executing code can be used to generate, store, and play generated media content. In the first block 1302 of method 1300, data associated with a first token is sent to a network address of a distributed ledger via a playback device and through the network. The address can be associated with a generated media smart contract configured to generate a generated media content model. In some examples, the first token can be an NFT, and optionally, multiple such tokens are sent to the smart contract address.

[0159] In block 1304, method 1300 receives a generated media content model from a network address associated with a smart contract for generated media via a playback device. For example, when the smart contract is executed, it generates a specific generated media content model that is at least partially based on data associated with a first token. This generated media content model can be provided to a user, which can then generate newly created media content based on the first token data. In addition to the first token data, the smart contract can also generate different outputs based on other input parameters (e.g., sensor data, playback device characteristic data, playback device state, user listening history data, etc.), as described elsewhere, provided that such data is provided to the smart contract address. Additionally, or instead, the generated media content model provided to the user can output different media content based on one or more other input parameters.

[0160] Next, in block 1306, the playback device generates media content based on at least a portion of the generated media content model. In some examples, this generation involves accessing a library of existing media segments stored in the playback device or other suitable storage location (e.g., a remote server, another device on the local network, etc.). The playback device can then arrange the existing media segments selected from the library for playback according to the generated media content model. In block 1308, method 1300 plays the generated media content through the playback device.

[0161] Various examples of generated media playback are described herein. As those skilled in the art will see, a wide variety of generated media modules, algorithms, inputs, sensor data, and playback device configurations are conceived and can be used in accordance with this technology.

[0162] IV. Conclusion The above discussion of playback devices, controller devices, playback zone configurations, and media content sources provides only a few examples of operating environments in which the functions and methods described below can be implemented. Other operating environments and configurations of media playback systems, playback devices, and network devices not expressly described herein may also be applicable and suitable for implementing the functions and methods.

[0163] The above description discloses various exemplary systems, methods, apparatus, and products, including, in particular, firmware and / or software running on hardware. It is understood that such examples are merely illustrative and should not be considered limiting. For example, any or all of the embodiments or components of the firmware, hardware, and / or software can be implemented exclusively in hardware, exclusively in software, exclusively in firmware, or in any combination of hardware, software, and / or firmware. Thus, the examples provided are not the only ways to implement such systems, methods, apparatus, and / or products.

[0164] Furthermore, any reference to “examples” in this specification means that certain features, structures, or characteristics described in relation to the examples may be included in at least one example or embodiment of the present invention. The appearance of this phrase in various parts of this specification does not necessarily refer to the same example, nor are they mutually exclusive or alternative examples. Thus, the examples described herein, as explicitly and implicitly understood by those skilled in the art, can be combined with other examples.

[0165] This specification primarily presents other symbolic representations that directly or indirectly resemble exemplary environments, systems, procedures, steps, logical blocks, processes, and the operation of network-coupled data processing devices. These descriptions and representations of processes are typically used by those skilled in the art to most effectively convey the nature of their work to others skilled in the art. Numerous specific details are provided to provide a complete understanding of this disclosure. However, it will be understood by those skilled in the art that certain examples of the art can be implemented without specific details. In other examples, well-known methods, procedures, components, and circuits are not described in detail to avoid unnecessarily obscuring the aspects of the examples. Therefore, the scope of this disclosure is defined not by the foregoing description of the examples but by the appended claims.

[0166] If any of the appended claims is read to cover purely software and / or firmware embodiments, then at least one of the elements in at least one example is expressly defined herein to include a tangible, non-temporary medium, such as memory, DVD, CD, or Blu-ray®, for storing the software and / or firmware.

[0167] The disclosed technology is illustrated, for example, by various examples described below. These various examples of embodiments of the disclosed technology are described as numbered examples (1, 2, 3, etc.) for convenience. These are provided as examples and do not limit the disclosed technology. Any of the dependent examples may be combined in any combination, or each may be included as an independent example. Other examples may be presented in a similar manner.

[0168] Example 1: A method comprising the steps of receiving input parameters in a coordinator device, transmitting the input parameters from the coordinator device to a plurality of playback devices, each having a generating media module internally, and transmitting timing data from the coordinator device to the plurality of playback devices so that the playback devices simultaneously play generated media content based at least partially on the input parameters.

[0169] Example 2: Any one of the examples herein, wherein a first and a second playback device each plays different generated audio content based at least partially on input parameters.

[0170] Example 3: Any one of the examples herein, in which the input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity)), user mood data).

[0171] Example 4: Timing data includes at least one of clock data or one or more synchronization signals, in any of the examples provided herein.

[0172] Example 5: Any one of the examples herein, further comprising the step of sending a signal from a coordinator device to at least one of a plurality of playback devices to change the generating media module of a playback device.

[0173] Example 6: One of the examples herein, wherein the generated media content comprises at least one of the generated audio content or generated visual content.

[0174] Example 7: A generating media module includes an algorithm that automatically generates a new media output based on an input that includes at least input parameters, one of the methods described herein.

[0175] Example 8: A device comprising a network interface, one or more processors, and a tangible, non-temporary, computer-readable medium for storing instructions causing the device to perform an operation when executed by the one or more processors, wherein the operation includes the steps of: receiving input parameters via the network interface; transmitting the input parameters via the network interface to a plurality of playback devices, each having a generating media module internally; and transmitting timing data via the network interface to the plurality of playback devices so that the playback devices simultaneously play generated media content based at least in part on the input parameters.

[0176] Example 9: One of the devices in the examples herein, wherein the first and second playback devices each play different generated audio content based at least partially on input parameters.

[0177] Example 10: One of the examples provided herein, in which the input parameters include one or more of the following devices: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., camera, light, temperature sensor, thermostat, presence detector, microphone); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity)), user mood data).

[0178] Example 11: A device, one of the examples herein, in which the timing data includes at least one of clock data or one or more synchronization signals.

[0179] Example 12: Any one of the examples herein, the operation further includes the step of sending a signal from a coordinator device to at least one of a plurality of playback devices via a network interface to change the generating media module of a playback device.

[0180] Example 13: One of the examples provided herein, wherein the generated media content includes at least one of the generated audio content or generated visual content.

[0181] Example 14: A generating media module is one of the examples provided herein, comprising an algorithm that automatically generates a new media output based on an input that includes at least input parameters.

[0182] Example 15: A tangible, non-temporary, computer-readable medium for storing instructions that cause a device to perform an action when executed by one or more processors of the device, wherein the action includes the steps of: receiving input parameters in a coordinator device; transmitting the input parameters from the coordinator device to a plurality of playback devices, each having a generating media module inside; and transmitting timing data from the coordinator device to the plurality of playback devices so that the playback devices simultaneously play generated media content based at least in part on the input parameters.

[0183] Example 16: A computer-readable medium of any of the examples herein, in which a first and a second playback device each plays different generated audio content based at least partially on input parameters.

[0184] Example 17: A computer-readable medium in any of the examples herein, in which the input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity)), user mood data).

[0185] Example 18: Timing data is a computer-readable medium in any of the examples herein, comprising clock data or at least one of one or more synchronization signals.

[0186] Example 19: A computer-readable medium according to any example herein, further comprising the step of sending a signal from a coordinator device to at least one of a plurality of playback devices to change the generating media module of a playback device.

[0187] Example 20: A computer-readable medium, one of the examples herein, in which the generated media content includes at least one of the generated audio content or generated visual content.

[0188] Example 21: A computer-readable media module, one of the examples herein, includes an algorithm that automatically generates a new media output based on an input that includes at least input parameters.

[0189] Example 22: A method comprising the steps of: receiving input parameters in a coordinator device; generating first and second media content streams via a generating media module of the coordinator device; transmitting the first media content stream to a first playback device via the coordinator device; and transmitting the second media content stream to a second playback device via the coordinator device such that the first and second media content streams are played simultaneously through the first and second playback devices.

[0190] Example 23: Any one of the examples herein further comprising the step of transmitting timing data from a coordinator device to a first and a second playback device, respectively.

[0191] Example 24: Timing data includes at least one of clock data or one or more synchronization signals, in any of the examples provided herein.

[0192] Example 25: Any one of the examples herein, wherein the first and second media content streams are different.

[0193] Example 26: Any one of the examples herein, in which the input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric data (heart rate, temperature, respiratory rate, brain activity)), user mood data).

[0194] Example 27: Any one of the examples herein, further comprising the step of modifying the generating media module of the coordinator device.

[0195] Example 28: Any one of the examples herein, wherein each of the first and second generated media content streams includes at least one of generated audio content or generated visual content.

[0196] Example 29: A generating media module includes an algorithm that automatically generates a new media output based on an input that includes at least input parameters, one of the methods described herein.

[0197] Example 30: A device comprising a network interface, a generating media module, one or more processors, and a tangible, non-temporary, computer-readable medium for storing instructions causing the device to perform an operation when executed by the one or more processors, wherein the operation includes the steps of receiving input parameters via the network interface, generating first and second media content streams via the generating media module, transmitting the first media content stream to a first playback device via the network interface, and transmitting the second media content stream to a second playback device via the network interface so that the first and second media content streams are played simultaneously through the first and second playback devices.

[0198] Example 31: Any one of the examples herein, whose operation further includes the step of transmitting timing data to each of the first and second playback devices via a network interface.

[0199] Example 32: A device in any of the examples herein, in which the timing data includes at least one of clock data or one or more synchronization signals.

[0200] Example 33: Any one of the examples herein, wherein the first and second media content streams are different.

[0201] Example 34: One of the examples provided herein, in which the input parameters include one or more of the following devices: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., camera, light, temperature sensor, thermostat, presence detector, microphone); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity)), user mood data).

[0202] Example 35: The operation further includes the step of modifying the generating media module, as described in any one of the examples of this specification for any device.

[0203] Example 36: One of the examples herein, each of the first and second generated media content streams, including at least one of generated audio content or generated visual content.

[0204] Example 37: A generating media module is one of the examples herein, comprising an algorithm that automatically generates a new media output based on an input that includes at least input parameters.

[0205] Example 38: A tangible, non-temporary, computer-readable medium for storing instructions causing a coordinator device to perform an operation when executed by one or more processors of the coordinator device, wherein the operation includes the steps of: receiving input parameters in the coordinator device; generating first and second media content streams via a generating media module of the coordinator device; transmitting the first media content stream to a first playback device via the coordinator device; and transmitting the second media content stream to a second playback device via the coordinator device so that the first and second media content streams are played simultaneously via the first and second playback devices.

[0206] Example 39: A computer-readable medium of any one of the examples herein, further comprising the step of transmitting timing data from a coordinator device to a first and a second playback device, respectively.

[0207] Example 40: Timing data is a computer-readable medium in any of the examples herein, comprising clock data or at least one of one or more synchronization signals.

[0208] Example 41: A computer-readable medium from any of the examples herein, wherein the first and second media content streams are different.

[0209] Example 42: A computer-readable medium in any of the examples herein, in which the input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity)), user mood data).

[0210] Example 43: A computer-readable medium, one of the examples herein, whose operation further includes the step of modifying the generating media module of a coordinator device.

[0211] Example 44: A computer-readable medium, one of the examples herein, in which each of the first and second generated media content streams includes at least one of generated audio content or generated visual content.

[0212] Example 45: A computer-readable media module, one of the examples herein, comprising an algorithm that automatically generates a new media output based on an input that includes at least input parameters.

[0213] Example 46: A playback device comprising one or more amplifiers configured to drive one or more audio transducers, one or more processors, and a data storage device having instructions that cause the playback device to perform an operation when executed by one or more processors, the operation comprising: receiving one or more first input parameters in the playback device; generating first media content via the playback device based at least partially on one or more first input parameters, the operation comprising accessing a library stored in the playback device which includes a plurality of existing media segments, and arranging a first selection of existing media segments from the library for playback based at least partially on one or more input parameters according to a generated media content model; and playing the first generated media content via one or more amplifiers.

[0214] Example 47: A playback device of any one of the examples herein, the operation comprising: receiving one or more second input parameters different from a first input parameter in a playback device; generating second media content via the playback device based at least partially on one or more second input parameters, wherein the second media content is different from the first media content, and the generating step includes accessing a library and arranging a second selection of existing media segments from the library for playback based at least partially on one or more second input parameters according to a generated media content model; and playing the second generated media content via one or more amplifiers.

[0215] Example 48: The playback device of claim 1, wherein the step of arranging a first selection of existing media segments from a library for playback includes the step of arranging two or more of the existing media segments at least partially in a time offset.

[0216] Example 49: A playback device in any of the examples herein, wherein the step of arranging a first selection of existing media segments from a library for playback includes the step of arranging two or more of the existing media segments so that they overlap at least partially in time.

[0217] Example 50: A playback device in any of the examples herein, in which the step of arranging a first selection of existing media segments from a library for playback includes the step of applying different equalization adjustments to different existing media segments.

[0218] Example 51: A playback device according to any example herein, wherein the step of arranging a first selection of existing media segments from a library or playback includes the step of applying time-varying gain levels to different existing media segments.

[0219] Example 52: A playback device according to any example herein, wherein the step of arranging a first selection of existing media segments from a library or playback includes the step of randomizing the starting point for playback of a particular existing media segment.

[0220] Example 53: A playback device according to any example herein, wherein the first generated media content and the second generated media content each contain new media content.

[0221] Example 54: A playback device according to any of the examples herein, wherein the first generated media content includes audio content, and the multiple existing media segments include multiple existing audio segments.

[0222] Example 55: A playback device according to any example herein, wherein the first generated media content includes audio-visual content, and the multiple existing media segments include multiple existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0223] Example 56: A playback device from any one of the examples herein, further comprising the steps of receiving additional existing media segments via a network interface and updating the library to include at least additional existing media segments.

[0224] Example 57: A playback device, any one of the examples herein, in which the first and second input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity, speech characteristics), user mood data).

[0225] Example 58: A method comprising the steps of: receiving one or more first input parameters in a playback device; generating first media content via the playback device based at least partially on one or more first input parameters, the method comprising: accessing a library stored in the playback device which includes a plurality of existing media segments; and arranging a first selection of existing media segments from the library for playback based at least partially on one or more input parameters according to a generated media content model; and playing the first generated media content via the playback device.

[0226] Example 59: Any one of the examples herein, a method comprising: receiving one or more second input parameters different from a first input parameter in a playback device; generating second media content via the playback device based at least partially on one or more second input parameters, wherein the second media content is different from the first media content, and the generating step includes accessing a library and arranging a second selection of existing media segments from the library for playback based at least partially on one or more second input parameters according to a generated media content model; and playing the second generated media content via the playback device.

[0227] Example 60: The step of arranging a first selection of existing media segments from a library for playback includes the step of arranging two or more of the existing media segments so as to be at least partially time-offset, one of the methods of the examples herein.

[0228] Example 61: The step of arranging a first selection of existing media segments from a library for playback includes the step of arranging two or more of the existing media segments so that they overlap at least partially in time, one of the methods of the examples herein.

[0229] Example 62: The step of arranging a first selection of existing media segments from a library for playback includes one of the examples herein, wherein the step of applying different equalization adjustments to different existing media segments is included.

[0230] Example 63: Any one of the examples herein, wherein the step of arranging a first selection of existing media segments from a library or playback includes the step of applying a time-varying gain level to different existing media segments.

[0231] Example 64: Any one of the examples herein, wherein the step of arranging a first selection of existing media segments from a library or playback includes the step of randomizing the starting point for playback of a particular existing media segment.

[0232] Example 65: Any one of the examples herein, wherein the first generated media content and the second generated media content each include new media content.

[0233] Example 66: Any one of the examples herein, wherein the first generated media content includes audio content and the multiple existing media segments include multiple existing audio segments.

[0234] Example 67: Any one of the examples herein, wherein the first generated media content includes audio-visual content, and the multiple existing media segments include multiple existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0235] Example 68: Any one of the examples herein further comprising the steps of receiving further existing media segments via a network interface and updating a library to include at least further existing media segments.

[0236] Example 69: Any one of the examples herein, wherein the first and second input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity, speech characteristics)), user mood data.

[0237] Example 70: A tangible, non-temporary, computer-readable medium for storing instructions causing a playback device to perform an operation when executed by one or more processors of the playback device, wherein the operation includes the steps of: receiving one or more first input parameters in the playback device; generating first media content via the playback device based at least partially on one or more first input parameters, the steps of: accessing a library stored in the playback device which includes a plurality of existing media segments; and arranging a first selection of existing media segments from the library for playback based at least partially on one or more input parameters according to a generated media content model; and playing the first generated media content via the playback device.

[0238] Example 71: A computer-readable medium according to any example herein, the operation comprising: receiving one or more second input parameters different from a first input parameter in a playback device; generating second media content via the playback device based at least partially on one or more second input parameters, wherein the second media content is different from the first media content, and the generating step includes accessing a library and arranging a second selection of existing media segments from the library for playback based at least partially on one or more second input parameters according to a generated media content model; and playing the second generated media content via one or more amplifiers.

[0239] Example 72: A computer-readable media in any of the examples herein, the step of arranging a first selection of existing media segments from a library for playback includes the step of arranging two or more of the existing media segments so as to be at least partially time-offset.

[0240] Example 73: A computer-readable medium in any of the examples herein, in which the step of arranging a first selection of existing media segments from a library for playback includes the step of arranging two or more of the existing media segments so that they overlap at least partially in time.

[0241] Example 74: A computer-readable media example herein in which the step of arranging a first selection of existing media segments from a library for playback includes the step of applying different equalization adjustments to different existing media segments.

[0242] Example 75: A computer-readable media example herein in which the step of arranging a first selection of existing media segments from a library for playback includes the step of applying a gain level that changes over time to different existing media segments.

[0243] Example 76: A computer-readable media example herein in which the step of arranging a first selection of existing media segments from a library for playback includes the step of randomizing the starting point for playback of a particular existing media segment.

[0244] Example 77: A computer-readable medium according to any example herein, wherein the first generated media content and the second generated media content each contain new media content.

[0245] Example 78: A computer-readable medium according to any of the examples herein, wherein the first generated media content includes audio content, and the multiple existing media segments include multiple existing audio segments.

[0246] Example 79: A computer-readable medium according to any of the examples herein, wherein the first generated media content includes audio-visual content, and the multiple existing media segments include multiple existing audio segments, existing visual media segments, or existing audio-visual media segments.

[0247] Example 80: A computer-readable media from any one of the examples herein, further comprising the steps of receiving additional existing media segments via a network interface and updating a library to include at least additional existing media segments.

[0248] Example 81: A computer-readable medium in any of the examples herein, wherein the first and second input parameters include one or more of the following: physiological sensor data (e.g., biometric sensors, wearable sensors (heart rate, body temperature, respiratory rate, electroencephalogram)); network device sensor data (e.g., cameras, lights, temperature sensors, thermostats, presence detectors, microphones); environmental data (e.g., weather, temperature, time / day / week / month); playback device capability data (e.g., number and type of transducers, output power); playback device status (e.g., device temperature, battery level, current audio playback, playback device location, whether the playback device is coupled with another playback device); or user data (e.g., user identification information, number of users present, user location, user history data, user preference data, user biometric authentication data (heart rate, temperature, respiratory rate, brain activity, speech characteristics)), user mood data.

[0249] Example 82: A system comprising a first playback device and a second playback device. The first playback device comprises a first network interface, one or more first processors, and a data storage device having instructions that cause the first playback device to perform an operation when executed by one or more processors, the operation comprising the steps of receiving one or more input parameters, generating media content based at least partially on one or more input parameters, wherein the generated media content comprises a first part and at least a second part, and the generating step includes accessing a library stored in the playback device that includes a plurality of existing media segments, and arranging a selection of existing media segments from the library for playback based at least partially on one or more input parameters according to a generated media content model, and the first network A system comprising the steps of transmitting a signal via an interface including a second portion of media content generated and corresponding timing information, and causing playback of a first portion of the generated media content, wherein the second playback device comprises a second network interface, one or more audio transducers, one or more second processors, and a data storage device having instructions that cause the second playback device to perform an operation when executed by one or more second processors, the operation comprising the steps of receiving a signal transmitted from the first playback device via the second network interface, and playing back the second portion of the generated media content via one or more transducers in accordance with the timing information and substantially synchronized with the playback of the first portion of the generated media content.

[0250] Example 83: A system of any example herein further comprising a network device, the network device comprising a third network interface, one or more processors, and a data storage device having instructions that cause a third playback device to perform an operation when executed by the one or more processors, the operation comprising the steps of receiving a request from a first playback device via the third network interface over a data network, and, in response to receiving the request, sending an updated library of existing media segments to the first playback device via the third network interface over a data network.

[0251] Example 84: A network device comprising one or more of the following systems as described herein: a remote server, another playback device, a mobile computing device, a laptop, or a tablet.

[0252] Example 85: A system comprising a first playback device and a second playback device, which are communicably coupled over a local area network, wherein the first playback device comprises one or more first processors, one or more first audio transducers, and a data storage device having instructions that cause the first playback device to perform an operation when executed by one or more first processors, the operation comprising the steps of receiving one or more input parameters, generating first media content based at least partially on one or more input parameters, accessing a first library stored in the first playback device which includes a plurality of existing media segments, and arranging a selection of existing media segments from the first library for playback based at least partially on one or more input parameters according to a first generated media content model, and playing the first generated media content via one or more first audio transducers, the second playback device comprising a second network interface.

[0253] A system comprising one or more second audio transducers, one or more second processors, and a data storage device having instructions for causing a second playback device to perform an operation when executed by one or more second processors, wherein the operation includes the steps of generating a second media content based at least partially on one or more input parameters, wherein the second generated media content is substantially identical to a first generated media content, and the generating step includes the steps of accessing a second library stored in a second playback device, which includes a plurality of existing media segments, and arranging a selection of existing media segments from the second library for playback based at least partially on one or more input parameters according to a second generated media content model; and playing the second generated media content via one or more second audio transducers in synchronization with the playback of the first generated media content via the first playback device.

[0254] Example 86: Any one of the examples herein, wherein the first generated media content model and the second generated media content model are substantially identical.

[0255] Example 87: Any one of the examples herein, wherein the first library and the second library are substantially identical.

[0256] Example 88: A method comprising the following steps: accessing blockchain data stored on a distributed ledger via a playback device; generating media content via a playback device based at least partially on the blockchain data; wherein the generating step includes: accessing a library containing a plurality of existing media segments stored in the playback device; and arranging a selection of existing media segments from the library for playback according to a generated media content model and at least partially based on the blockchain data; and further, playing the generated media content via the playback device.

[0257] Example 89: Any one of the examples herein, wherein the NFT data includes one or more existing media segments, and accessing the NFT data includes storing one or more existing media segments in a library.

[0258] Example 90: A method of any one example herein, wherein blockchain data includes first non-fungible token (NFT) data, where the distributed ledger is the first distributed ledger, and the method further includes accessing data related to a second NFT stored in a second distributed ledger via a playback device, where arranging a selection of existing media segments from a library for playback according to a generated media content model is at least partially based on both the first NFT data and the second NFT data.

[0259] Example 91: Any one method of this specification wherein the first distributed ledger is associated with a first blockchain layer, and the second distributed ledger is associated with a second blockchain layer different from the first blockchain layer.

[0260] Example 92: One of the examples herein in which blockchain data is associated with a playlist.

[0261] Example 93: Any one of the examples herein, wherein blockchain data relies, at least in part, on transactions recorded on a distributed ledger, including non-fungible tokens (NFTs).

[0262] Example 94: Arranging a selection of existing media segments from a library according to a generative media content model is one of the methods in the examples herein, which is at least partially based on one or more input parameters.

[0263] Example 95: Any one of the examples herein, wherein the input parameters include one or more of the following: physiological sensor data, networked device sensor data, environmental data, regenerative device characteristic data, regenerative device status, user listening history data, oracle data stored via a distributed ledger, or user data.

[0264] Example 96: Any one of the examples herein in which user listening history data is stored via a distributed ledger.

[0265] Example 97: Any one of the examples herein, in which accessing blockchain data includes connecting to a user wallet that holds non-fungible tokens (NFTs).

[0266] Example 98: Any one of the examples herein, in which accessing blockchain data includes accessing a code associated with a physical medium object via a control device (e.g., a QR code or other code engraved on custom vinyl or other medium).

[0267] Example 99: Arranging a selection of existing media segments from a library for playback includes any one of the examples herein, which involves arranging two or more existing media segments in a manner that is at least partially time-offset.

[0268] Example 100: Arranging a selection of existing media segments from a library for playback includes any one of the examples herein, which involves arranging two or more of the existing media segments so that they overlap at least partially in time.

[0269] Example 101: A method comprising the following steps: sending data associated with a first token to a network address of a distributed ledger via a playback device via a network, the address being associated with a generating media smart contract configured to generate a generated media content model; receiving the generated media content model from the network address associated with the generating media smart contract via a playback device; and generating media content via a playback device, at least partially based on the generated media content model, wherein the generating step includes: accessing a library containing a plurality of existing media segments; and arranging a selection of existing media segments from the library for playback according to the generated media content model, the method further comprising playing the generated media content via a playback device.

[0270] Example 102: The token data includes first non-fungible token (NFT) data, and the method further includes transmitting, via a playback device, data related to a second NFT stored in a distributed ledger to a network address associated with a generated media smart contract; receiving, via the playback device, a second generated media content model from a network address associated with the generated media smart contract, where the second generated media content model is different from the first generated media content model; generating, via the playback device, second media content at least partially based on the second generated media content model; and playing, via the playback device, the second generated media content; The method according to any one of the examples herein.

[0271] Example 103: The method according to any one of the examples herein, wherein the token data is associated with a curated playlist.

[0272] Example 104: The method according to any one of the examples herein, wherein the token data depends at least in part on a transaction recorded in a distributed ledger containing tokens.

[0273] Example 105: The step of arranging the selection of existing media segments from a library according to a generated media content model is further based at least in part on one or more input parameters. The method according to any one of the examples herein

[0274] Example 106: The method according to any one of the examples herein, wherein the input parameters include one or more of physiological sensor data, networked device sensor data, environmental data, playback device characteristic data, playback device state, user listening history data, oracle data stored via a distributed ledger, or user data.

[0275] Example 107: The method according to any one of the examples herein, wherein the user's listening history data is stored via a distributed ledger.

[0276] Example 108: Any one of the methods of the present specification, further comprising accessing token data by connecting to a user wallet that stores the token data before transmitting the token data.

[0277] Example 109: Any one of the methods of the present specification, further comprising accessing token data via a code (e.g., a QR code or other code printed on custom vinyl or other media) associated with a physical media object via a control device before transmitting the token data.

[0278] Example 110: Any one of the methods of the present specification, wherein the step of arranging a selection of existing media segments from a library for playback includes arranging two or more of the existing media segments in at least a partially temporally offset manner.

[0279] Example 111: Any one of the methods of the present specification, wherein the step of arranging a selection of existing media segments from a library for playback includes arranging two or more of the existing media segments in at least a partially temporally overlapping manner.

[0280] Example 112: Any one of the methods of the present specification, wherein the first generated media content and the second generated media content each include new media content.

[0281] Example 113: One or more tangible, non-transitory computer-readable media storing instructions executable by one or more processors to cause a media playback system or playback device to perform operations according to any one of the methods of the present specification.

[0282] Example 114: A media playback system comprising one or more processors; and a computer-readable medium holding any one of the methods of the present specification.

[0283] Example 115: A playback device comprising one or more processors and a computer-readable medium holding any one of the methods of the examples herein.

Claims

1. In a coordinator device, the steps include receiving input parameters, The steps include: transmitting the input parameters from the coordinator device to multiple playback devices, each having an internal generation media module; The steps include: transmitting timing data from the coordinator device to the plurality of playback devices so that the playback devices simultaneously play generated media content based at least partially on the input parameters; A method for providing this.

2. The method according to claim 1, wherein a first and a second playback device each plays different generated audio content based at least partially on input parameters.

3. In a coordinator device, the steps include receiving input parameters, The steps include generating first and second media content streams via the generation media module of the coordinator device, The steps include transmitting the first media content stream to the first playback device via the coordinator device, The steps include: transmitting the second media content stream to the second playback device via the coordinator device so that the first and second media content streams are played simultaneously through the first and second playback devices; A method for providing this.

4. The method according to claim 3, further comprising the step of transmitting timing data from the coordinator device to each of the first and second playback devices.

5. The method according to claim 4, wherein the first and second media content streams are different.

6. The method according to claim 4, wherein the timing data includes at least clock data or one or more synchronization signals.

7. A method according to any one of claims 3 to 6, wherein each of the first and second generated media content streams includes at least one of generated audio content or generated visual content.

8. A method according to any one of claims 3 to 7, wherein the generating media module includes an algorithm for automatically generating a new media output based on an input that includes at least the input parameters.

9. The steps include receiving one or more first input parameters in a playback device, The steps include generating a first media content via the playback device based at least partially on one or more of the first input parameters, The step of generating includes, The steps include: accessing a library stored in the playback device that includes multiple existing media segments; The steps include: arranging a first selection of existing media segments from the library for playback based at least partially on one or more input parameters according to a generated media content model; The steps include playing the first media content generated via the playback device, A method that includes this.

10. The method according to claim 9, further, The playback device includes the step of receiving one or more second input parameters that are different from the first input parameter, A step of generating a second media content via the playback device, at least partially based on one or more of the second input parameters, wherein the second media content is different from the first media content. Equipped with, The above generation step is, The steps include: accessing the aforementioned library, The steps include arranging a second selection of existing media segments from the library for playback based at least partially on one or more of the second input parameters according to the generated media content model, The steps include playing the generated second media content via the playback device, Includes.

11. The method according to claim 10, wherein the first media content generated and the second media content generated each include new media content.

12. The method according to any one of claims 9 to 11, wherein the generated first media content includes audio content, and the plurality of existing media segments include a plurality of existing audio segments.

13. The method according to any one of claims 9 to 12, wherein the generated first media content includes audio-visual content, and the plurality of existing media segments include a plurality of existing audio segments, an existing visual media segment, or an existing audiovisual media segment.

14. The steps include receiving further existing media segments via a network interface, The steps include updating the library to include at least the further existing media segments, The method according to any one of claims 9 to 13, further comprising:

15. It is a playback device, One or more amplifiers configured to drive one or more audio transducers, One or more processors, A data storage device having an instruction that causes the playback device to perform the method described in any one of claims 9 to 14 when executed by one or more processors, A playback device equipped with the following features.

16. A first playback device based on the playback device described in claim 15, A second network interface, a second playback device having one or more audio transducers, and one or more second processors, A system equipped with, The second processor is, The system receives signals transmitted from the first playback device via a second network interface, which have a second portion of the generated media content and timing information. The second portion of the generated media content is played back in accordance with the timing information, so as to be substantially synchronized with the playback of the first portion of the generated media content via the first playback device, via one or more transducers. system.

17. The system according to claim 16, further comprising a network device, the network device having the following: The third network interface; One or more processors; and A data storage device that stores instructions that can be executed by one or more processors. When the instruction is executed by one or more processors, the third playback device: The data network receives a request from the first playback device via a third network interface. In response to receiving the aforementioned request, the updated library for existing media segments is sent to the first playback device via the third network interface of the data network.

18. The system according to claim 17, wherein the network device comprises one or more remote servers, other playback devices, mobile computing devices, laptops, or tablets.

19. A first playback device based on the playback device described in claim 15, A second playback device having a second network interface, one or more second audio transducers, one or more second processors, and a data storage device storing instructions executable by one or more second processors, A system equipped with, When the instruction is executed by one or more second processors, the second playback device: By accessing a second library stored on a second playback device containing multiple existing media segments, and According to the second generated media content model, by arranging a selection of existing media segments from a second library for playback based at least partially on one or more input parameters, Generate a second media content based at least partially on one or more input parameters: and The generated second media content is played in synchronization with the playback of the generated first media content via the first playback device, via one or more second audio transducers. system.

20. The system according to claim 19, wherein the first generated media content model and the second generated media content model are substantially identical.

21. The system according to claim 19 or 20, wherein the first library and the second library are substantially identical.

22. The method according to claims 1 to 14, wherein the method involves receiving one or more input parameters. Blockchain stored on a distributed ledger via a playback device This includes accessing blockchain data, where one or more parameters consist of blockchain data.

23. A method according to claim 22, wherein the blockchain data consists of one or more existing media segments, and accessing the blockchain data includes storing one or more existing media segments in a library.

24. A method according to claim 22 or 23, wherein the blockchain data comprises first non-fungible token (NFT) data, wherein the distributed ledger is the first distributed ledger, and further the method comprises accessing data relating to a second NFT stored in a second distributed ledger via a playback device, and further wherein the arrangement of a selection of existing media segments from a library for playback according to a generated media content model is at least partially based on both the first NFT data and the second NFT data.

25. A method according to any one of claims 22 to 24, wherein a first distributed ledger is associated with a first blockchain layer, and a second distributed ledger is associated with a second blockchain layer different from the first blockchain layer.

26. A method according to any one of claims 22 to 25, wherein blockchain data is associated with a playlist.

27. A method according to any one of claims 22 to 25, wherein at least a portion of the blockchain data relies on transactions recorded on a distributed ledger, which includes non-fungible tokens (NFTs).

28. A method according to any one of claims 22 to 25, wherein the selection of existing media segments from a library according to a generated media content model is further at least in part on one or more additional input parameters.

29. A method according to any one of claims 22 to 28, wherein access to blockchain data is made by connecting to a user wallet holding non-fungible tokens (NFTs).

30. A method according to any one of claims 22 to 29, wherein access to blockchain data is made by accessing code associated with a physical media object via a control device.

31. The playback device sends data associated with a first token to a network address on a distributed ledger via the network, where this address is associated with a generated media smart contract configured to generate a generated media content model; The playback device receives the generated media content model from the network address associated with the generated media smart contract; The process includes the step of generating media content using a playback device based at least partially on the generated media content model. The above generation step is, Accessing a library consisting of multiple existing media segments; It has the ability to arrange a selection of existing media segments from the library for playback according to the generated media content model; Furthermore, the playback device includes a step of playing the generated media content. method.

32. The method according to claim 31, wherein the token data comprises first non-fungible token (NFT) data, Furthermore, this method, The steps include: sending data related to a second NFT stored in a distributed ledger via a playback device to a network address associated with the generating media smart contract; The playback device receives a second generated media content model from a network address associated with the generated media smart contract, wherein the second generated media content model is different from the first generated media content model; The steps include: generating a second media content based at least partially on a second generated media content model using a playback device; The steps include: playing the second generated media content using the playback device; It is equipped with.

33. The method according to claim 31 or 32, wherein the token data is associated with a curated playlist.

34. A method according to any one of claims 31 to 33, wherein the token data relies at least in part on transactions recorded in a distributed ledger in which the tokens are involved.

35. A method according to any one of claims 31 to 34, wherein the selection of existing media segments from a library according to a generated media content model is further based at least in part on one or more input parameters.

36. A method according to any one of claims 31 to 35, further comprising the step of connecting to a user wallet that stores token data and accessing the token data before transmitting the token data.

37. A method according to any one of claims 31 to 36, further comprising the step of accessing the token data via a code associated with a physical medium object via a control device before transmitting the token data.

38. A method according to any one of claims 31 to 36, wherein the first and / or second input parameters consist of one or more of the following: Physiological sensor data Network device sensor data Environmental data Playback device capability data Playback device status User data Oracle data stored via a distributed ledger.

39. A method according to any one of claims 22 to 38, wherein arranging a first selection of existing media segments from a library for playback includes at least one of the following: Placing two or more existing media segments at least partially offset in time; Arranging two or more existing media segments so that they overlap at least partially in time; Applying different equalization adjustments to different existing media segments; Applying a time-varying gain level to different existing media segments; Randomizing the playback start point of a specific existing media segment.

40. A method according to any one of claims 22 to 39, wherein the user's viewing history data is stored via a distributed ledger.

41. The method according to any one of claims 22 to 40, wherein the first generated media content and the second generated media content each consist of novel media content.

42. One or more tangible, non-transient, computer-readable media that store instructions executable by one or more processors, causing a media playback system or playback device to perform the operations described in any one of claims 1 to 14 or 22 to 41.

43. It is a media playback system. One or more processors and The computer-readable medium according to claim 42.

44. A playback device One or more processors and The computer-readable medium according to claim 42.

Citation Information

Patent Citations

  • Content contract system, content contract method, right holder terminal, assignee terminal, control terminal, content storage server, right holder program, assignee program, control program, and content storage program

    JP2020068388A

  • Distributed digital content distribution systems and processes using blockchain

    JP2020521257A

  • Improving computer system security using a biometric authentication gateway for user service access using split and distributed secret encryption keys

    JP2022525765A

  • Music content generation

    WO2021163377A1